Intelligent voice interaction and analysis system and method based on large model
Through the intelligent voice interaction system of multimodal input and deep learning model combined with the knowledge base module, the existing system's recognition accuracy and semantic understanding problems in complex environments are solved, efficient voice command analysis and multi-round dialogue management are realized, and the system's adaptability and user interaction are improved.
Patent Information
- Application Number
- CN202510415682.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-29
AI Technical Summary
The existing voice interaction systems have low recognition accuracy, insufficient semantic understanding ability, poor adaptability and lack of user interaction in complex environments, making it difficult to deal with dynamic scenarios and real-time user needs.
It adopts multi-modal input module, real-time feedback mechanism, multi-module collaborative architecture and interactive adjustment interface, combined with deep learning model and knowledge base module, realizes efficient processing of voice signals and multi-round dialogue management, supports multi-language and dialect, performs speech preprocessing, semantic understanding and result optimization, and provides a variety of output formats.
It improves the accuracy of speech recognition, can deeply analyze the intention and context information of voice commands, enhances the adaptability and flexibility of the system, supports a variety of input and output methods, and meets users' personalized needs.
Smart Images

Figure CN120388563A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically to an intelligent voice interaction and analysis system and method based on a large model. Background Art
[0002] In today's digital age, the demand for voice interaction is increasing day by day, and it is widely used in many fields such as intelligent customer service, smart home, and in-vehicle voice interaction. Traditional voice interaction systems mainly rely on simple speech recognition technology and preset instruction sets. These methods have problems such as low efficiency, insufficient semantic understanding ability, and difficulty in processing complex voice instructions. In recent years, with the development of artificial intelligence and natural language processing technologies, speech interaction technologies based on deep learning have gradually emerged. However, the existing technologies still have the following limitations:
[0003] Limited recognition accuracy: The speech recognition accuracy of existing systems is relatively low in complex environments (such as high noise, multiple people talking), making it difficult to meet high-precision requirements.
[0004] Insufficient semantic understanding ability: The semantic understanding of voice instructions mainly focuses on the surface meaning, and it is difficult to deeply analyze the user's intention and context information.
[0005] Poor adaptability: The voice interaction ability of existing systems in different fields is limited, and it is difficult to quickly adapt to the needs of new fields.
[0006] Lack of user interaction: Most existing systems are one-way output, lacking real-time interaction with users. Users cannot adjust and optimize the recognition results according to their needs in real time.
[0007] To solve the above problems, some voice interaction and analysis tools based on artificial intelligence technology have emerged in recent years. These tools usually adopt traditional speech recognition technology or simple deep learning models, and can realize basic functions such as speech-to-text conversion, instruction recognition, and simple semantic understanding. For example, some tools use hidden Markov models or recurrent neural networks to process voice signals, and can recognize simple instructions or common problems in speech. However, the existing technologies still have the following limitations:
[0008] Difficulty in handling dynamic scenarios: Existing tools often have difficulty effectively capturing and analyzing the user's real-time intention and emotional changes when dealing with dynamic voice interaction. For example, in the intelligent customer service scenario, the recognition and response to the user's emotions.
[0009] Lack of user interactivity: Most existing voice interaction systems are one-way output, lacking real-time interaction with users. Users cannot adjust and optimize the recognition results according to their needs in real time, resulting in insufficient flexibility and practicality of the system. Summary of the Invention
[0010] The object of the present invention is to provide an intelligent voice interaction and analysis system and method based on a large model to solve the problems raised in the above background technology.
[0011] To achieve the above object, the present invention provides the following technical solution: An intelligent voice interaction and analysis system based on a large model, comprising:
[0012] Multimodal input module: Supports hybrid interaction of voice input and text input, is compatible with multiple languages and dialects, and allows users to supplement and correct instructions by voice or text;
[0013] Real-time feedback mechanism: Immediately starts recognition and analysis after receiving input, dynamically displays the processing progress and preliminary results, and supports users to adjust requirements in real time;
[0014] Multi-module collaborative architecture: Integrates voice preprocessing, large model core, result optimization, and knowledge base modules to achieve end-to-end processing from voice signals to executable instructions;
[0015] Result export function: Supports exporting speech-to-text, semantic parsing results, and operation feedback as text, audio, or structured data files;
[0016] Interactive adjustment interface: Provides a graphical operation interface for users to adjust instruction details, optimize answer logic, and select output formats.
[0017] Preferably, the voice preprocessing module further includes:
[0018] Signal cleaning unit: Uses an adaptive filtering algorithm to eliminate background noise and unifies the sound pressure level of audio segments through volume normalization technology;
[0019] Feature extraction unit: Extracts feature parameters such as Mel-frequency cepstral coefficients MFCC, fundamental frequency period, and speech rate variation coefficient to construct a high-dimensional representation of the voice signal;
[0020] Lightweight semantic understanding engine: Coarse-grained classification of voice instructions based on a pre-trained language model: query / instruction / dialogue, and extracts key entities such as time, place, and object.
[0021] Preferably, the large model core module includes the following innovations:
[0022] Multi-granularity speech recognition model: Adopts a Transformer architecture to fuse the acoustic model and the language model, and supports the recognition of dialect variants of more than 8 languages including Mandarin, Cantonese, and English;
[0023] Context-aware parser: Achieves multi-turn dialogue context tracking through the attention mechanism and automatically supplements omitted information;
[0024] Domain Adaptation Fine-tuning Framework: It supports loading knowledge distillation parameters in professional fields such as medicine and finance, and adjusts the model weights in real time based on user feedback.
[0025] Preferably, the result optimization module has the following functions:
[0026] Semantic Refinement Algorithm: It uses the BERT model to eliminate redundancy and reorganize the logic of the generated text to ensure that the answer conforms to the expression specification of "presenting the conclusion first - hierarchical argumentation";
[0027] Multi-modal Output Adaptation Layer: It automatically converts the format according to the output scenario, such as generating technical documents in Markdown format, device control instructions in JSON structure, or emotional speech synthesis waveforms;
[0028] Confidence Screening Mechanism: It sorts the parsing results by confidence, only presents the information above the threshold, and marks the possible error types of the low-confidence content.
[0029] Preferably, the knowledge base module achieves the following technical effects:
[0030] Hybrid Knowledge Graph: It constructs a three-layer knowledge architecture including general encyclopedias, domain term libraries, and scenario templates, where the smart home template contains the triple relationship of "device - action - environment";
[0031] Incremental Update Mechanism: It automatically discovers new knowledge by monitoring user conversation logs, and uses the graph neural network GNN to achieve the dynamic evolution of knowledge embedding vectors;
[0032] Template-based Response Generation: It pre-sets more than 50 scenario response templates and supports users to customize parameter mapping rules.
[0033] A method for an intelligent voice interaction and analysis system based on a large model includes the following steps:
[0034] Receive questions or instructions input by the user through voice or text;
[0035] Preprocess the input voice, including noise reduction, volume adjustment, feature extraction, and preliminary semantic understanding, extract key features such as Mel Frequency Cepstral Coefficients (MFCC), intonation, and speech rate in the voice signal, and identify the intention category and key information of the voice instruction;
[0036] Use a deep learning large model to recognize and analyze the preprocessed voice. Based on recurrent neural networks (RNN), Transformer architectures, and other efficient and accurate network architectures, convert the voice signal into text, parse the user's intention and context information, and generate a preliminary answer or operation instruction;
[0037] Optimize the preliminary results generated by the large model, including semantic optimization to ensure coherent content and clear logic, screen important information according to user needs, remove redundant content, and typeset and format the results according to the output format specified by the user;
[0038] According to the input requirements of the user, call relevant professional knowledge and templates from the knowledge base covering multiple fields such as intelligent customer service, smart home, and in-vehicle voice interaction to provide background information support for the large model and assist in generating the final result.
[0039] Preferably, it includes the following steps:
[0040] The user performs voice input by clicking the "Voice Input" button in the interface or directly enters specific instructions or questions in the text box, clarifies the interaction goal, and then clicks the "Send" button after completion;
[0041] Clean the voice input by the user, remove background noise through a filtering algorithm, automatically detect and adjust the voice volume using volume normalization technology to enhance voice recognizability, then extract the key features in the voice signal, and classify the intent of the voice command and extract key information;
[0042] Input the preprocessed voice into the deep learning large model, which performs speech recognition to convert the voice signal into text, performs semantic parsing to accurately understand the user's intent and context information, supports multi-turn dialogue management, and generates a preliminary answer or operation instruction;
[0043] Optimize the preliminary results, perform in-depth semantic optimization on the answer or operation instruction, screen key information according to user needs, remove unnecessary redundant information, and perform typesetting adjustment according to the voice answer, text display, and operation feedback output formats specified by the user;
[0044] The system calls relevant knowledge from the knowledge base storing multi-domain knowledge bases and professional term bases according to the user's input requirements to provide professional and accurate support for result generation, and the knowledge base is updated in real time to maintain timeliness.
[0045] Preferably, it includes the following steps:
[0046] Receive the user's requirements input in various languages or dialects in the form of voice or text;
[0047] Preprocess the input voice, first perform noise reduction and volume adjustment to improve the voice quality, then extract key features such as MFCC, intonation, and speech rate, and at the same time perform semantic understanding on the voice command to determine the intent category such as query, instruction, or dialogue and extract key information;
[0048] Use a large deep learning model to recognize and analyze the preprocessed speech, convert the speech into text, parse the user's intention and context information, manage multi-round conversations, and generate preliminary results including speech-to-text conversion, semantic parsing results, answers, or operation feedback;
[0049] Optimize the content of the preliminary results to ensure that the answers are logically rigorous and well-organized, remove redundant information, screen important information according to the user's needs, and format and adjust the results according to the user-specified format;
[0050] Leverage the knowledge base to provide domain knowledge support. The knowledge base stores knowledge and predefined templates in multiple fields such as intelligent customer service, smart home, and in-vehicle voice interaction, calls relevant knowledge according to the user input, and the knowledge base is updated in real time based on user feedback and system operation conditions.
[0051] Preferably, it includes the following steps:
[0052] The user inputs questions or instructions to the system in the form of speech or text, and the system supports multiple languages and dialects;
[0053] Perform a series of preprocessing operations on the input speech, including noise reduction, volume adjustment, feature extraction, and semantic understanding, to obtain the key features of the speech, the instruction intention, and key information;
[0054] Input the preprocessed speech into a large deep learning model based on multiple advanced network architectures. The model performs speech recognition and semantic parsing, understands the user's intention and context, generates preliminary answers or operation instructions, and supports multi-round conversations;
[0055] Optimize the preliminary results, perform semantic optimization to make the content coherent and clear, screen key information according to the user's needs, remove redundancy, and finely format the results according to the user-specified output format;
[0056] Use the knowledge base to provide assistance for result generation. The knowledge base covers knowledge and templates in multiple fields, calls relevant knowledge according to the user input, and is updated in real time to ensure the accuracy and professionalism of the results. At the same time, it supports users to create and save custom templates according to their needs.
[0057] Preferably, it includes the following steps:
[0058] Receive the user's needs input in the form of speech or text, and support multiple languages and dialects;
[0059] Preprocess the input speech, including noise reduction, volume adjustment, extraction of MFCC, intonation, speech rate key features, and preliminary semantic understanding, to determine the intention category and key information of the speech instruction;
[0060] Use a large deep learning model to recognize and analyze pre-processed speech, convert speech into text, analyze user intent and context, manage multiple rounds of dialogue, and generate preliminary responses or operational instructions.
[0061] Optimize the preliminary results, perform semantic optimization on the content, filter important information according to user needs, remove redundant content, and adjust the layout according to the user-specified voice response and text display operation feedback format;
[0062] With the help of the knowledge base, domain knowledge support is provided. The knowledge base stores multi-domain knowledge and predefined templates, and calls relevant knowledge based on user input. The knowledge base is updated in real time based on user feedback and system operation status. The system also has a model fine-tuning function, which can perform targeted fine-tuning on large models based on domain requirements and feedback input by users to improve model adaptability and accuracy.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] The large-scale model-based intelligent voice interaction and analysis system and method proposed in this paper effectively improves speech recognition accuracy by introducing advanced deep learning models, particularly pre-trained models based on recurrent neural networks (RNNs) and Transformer architectures. During the pre-processing phase, the system performs noise reduction, volume adjustment, and feature enhancement on the speech, further optimizing the input data quality and providing the model with clearer and more accurate speech information.
[0065] Through deep learning models and contextual understanding technology, voice commands can be analyzed in depth across multiple dimensions. The system can not only identify specific commands in the voice, but also analyze the user's intent, contextual information, and emotional tendencies.
[0066] The integration of a knowledge base module and model fine-tuning technology significantly enhances the system's adaptability. The knowledge base module stores a wealth of domain knowledge and templates, providing strong context for the model. Furthermore, the system automatically loads the corresponding domain datasets and parameters based on user-entered domain requirements, fine-tuning the model.
[0067] An interactive user interface allows users to ask questions or provide commands via voice or text input and view responses and feedback in real time. The system allows users to modify, adjust, or supplement results, allowing them to optimize interactive results based on their needs. The system responds to user actions in real time, rapidly updating and re-displaying results. This interactive design not only increases user engagement but also enables the system to dynamically optimize results based on user feedback, meeting personalized needs.
[0068] Supports multiple input methods, including voice input and text input, to meet the needs of different users. At the same time, the system supports multiple output formats, such as voice answers, text displays, operation feedback, etc., and users can choose the appropriate output method according to their own needs. This multi-modal input and output ability makes the system more flexible and better able to integrate into different application scenarios and work processes. Brief Description of the Drawings
[0069] Figure 1 It is a block diagram of the system of the present invention;
[0070] Figure 2 It is a flow chart of intelligent voice interaction and analysis of the present invention. Detailed Description of the Embodiments
[0071] In order to clearly and completely describe the objectives, technical solutions of the present invention, and make the advantages more clearly understood, the following further details the embodiments of the present invention with reference to the drawings. It should be understood that the specific embodiments described herein are some embodiments of the present invention, rather than all embodiments, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0072] Embodiment 1, please refer to Figures 1 to 2 , the present invention provides a technical solution: an intelligent voice interaction and analysis system based on a large model, including:
[0073] Voice Interaction Module
[0074] Multi-modal Input Support: Users can ask questions or give instructions through voice input, and the system supports multiple languages and dialects. In addition, users can also supplement or correct voice instructions through text input.
[0075] Real-time Feedback and Interaction: After receiving user input, the system starts the recognition and analysis process in real time, and dynamically displays the progress and preliminary results in the form of voice or text. Users can view the results in real time and adjust or supplement their needs through voice or text.
[0076] Result Display and Export: Displays the results of voice interaction in an intuitive manner, including speech-to-text, semantic parsing results, answers or operation feedback, etc. Supports exporting the results in multiple formats, such as text files, audio files, etc., for convenient subsequent use by users.
[0077] Voice Preprocessing Module
[0078] Speech Cleaning: First, noise reduction is performed, using filtering algorithms to remove background noise from the speech signal and improve speech clarity. Next, volume adjustment is performed, automatically detecting and adjusting the speech volume to enhance speech intelligibility. Volume normalization technology ensures consistent volume levels across all audio clips, preventing volume discrepancies from affecting subsequent processing.
[0079] Feature extraction: Extract key features from speech signals, such as Mel-Frequency Cepstral Coefficients (MFCCs), intonation, and speaking rate, for subsequent speech recognition and semantic analysis. MFCCs are widely used speech features that effectively capture the spectral characteristics of speech. Intonation and speaking rate can help the system better understand the user's intent and emotional state, supporting semantic understanding by analyzing the frequency variations and duration of speech signals.
[0080] Semantic Understanding: The system categorizes voice commands into different categories, such as queries, instructions, or conversations. For example, when a user asks, "What's the weather like tomorrow?" the system recognizes it as a query. After determining intent, the system further extracts key information, such as time, location, or object. This information serves as the foundation for in-depth analysis, helping the system generate accurate responses or perform corresponding actions.
[0081] Large model core module
[0082] Speech Recognition: Speech recognition technology is a core feature of modern intelligent interaction. Leveraging advanced deep learning models, it efficiently converts speech signals into text. This technology not only supports multiple languages but also accurately recognizes various dialects, ensuring accurate speech transcription services for users in diverse language environments. It easily handles Mandarin, Cantonese, Minnan dialect, English, French, German, and other languages, greatly expanding the scope of voice interaction.
[0083] Semantic parsing: Based on voice recognition, the semantic parsing function further enhances the intelligent level of interaction. It can deeply understand the semantics of voice commands and accurately parse the user's intentions and contextual information. For example, when a user says "What's the weather like tomorrow?", the system not only understands the core concept of "weather", but also recognizes the time dimension of "tomorrow", thereby providing accurate weather forecast information. In addition, semantic parsing also supports multi-round dialogue management, which can continuously optimize the dialogue process based on contextual information, making the interaction more natural and smooth. For example, in a conversation, the user first asks "What's the weather like tomorrow?" and then asks "What about the day after tomorrow?" The system can automatically recognize and continue the context, and directly provide the weather information for the day after tomorrow without the user having to repeat it.
[0084] Content Analysis: Content analysis is a crucial step in transforming the results of semantic parsing into actual operations. Based on the user's voice commands, the system can generate corresponding responses or operation instructions. Taking the smart home scenario as an example, when the user issues the command "Turn on the living room light", the system quickly parses the command content, identifies the device object "living room light" and the operation instruction "turn on", and through the connection with the smart home system, accurately executes the operation of "turning on the living room light". This process is not only efficient but also ensures the accuracy of the operation, allowing users to easily control home devices with simple voice commands.
[0085] Model Fine-tuning: To further improve the adaptability and accuracy of the model, the system also has a powerful model fine-tuning function. According to the domain requirements input by the user, the system can perform targeted fine-tuning on the large model. For example, in the medical field, by loading specific medical datasets and parameters, the model can accurately understand medical terms and related instructions; in the financial field, by loading financial data, the model's understanding of financial terms and business processes can be optimized. In addition, model fine-tuning supports personalized adjustments. Based on the user's feedback and annotation information, the system can dynamically adjust the model parameters and continuously optimize the recognition and analysis results. For example, if the user is not satisfied with the recognition result of a certain voice command, they can provide feedback through annotation, and the system will automatically adjust the model parameters according to this feedback information, so as to provide more accurate results in subsequent interactions. This dynamic optimization mechanism ensures that the model can continuously adapt to the user's needs and provide more personalized services.
[0086] Result Optimization Module
[0087] Content Optimization: When generating responses or operation instructions, in-depth semantic optimization of the content will be carried out. Ensure that the logic of the response is rigorous and well-organized, enabling users to understand smoothly. At the same time, all unnecessary redundant information will be carefully checked and removed, presenting the core points in a concise and clear manner, further optimizing the expression of the response, and making the language more accurate and natural.
[0088] Result Screening: According to the user's clear needs, the most relevant and important parts are accurately screened out from a large amount of information. Remove the redundant content that has nothing to do with the user's goals, highlight the key results, and ensure that users can quickly obtain the information they really need, improving the efficiency and value of information acquisition.
[0089] Format Adjustment: According to the output format specified by the user, fine typesetting optimization of the results is carried out. Whether it is voice responses or various formats such as text displays, it can be flexibly adapted to meet the user's usage needs in different scenarios and enhance the user experience.
[0090] Knowledge Base Module
[0091] Domain Knowledge Storage: We have a multi-domain knowledge base, covering expertise and datasets in areas such as intelligent customer service, smart home, and in-vehicle voice interaction, providing rich background information for the model. We also store a library of specialized terminology to ensure the professionalism and accuracy of generated results.
[0092] Template Management: Stores predefined voice interaction templates, allowing users to quickly select the appropriate template based on their needs. This allows users to create and save custom templates based on their needs, increasing system flexibility.
[0093] Real-time updates: Knowledge base content is updated in real time based on user feedback and system performance to optimize system performance. The latest domain knowledge and data are regularly synchronized from external data sources to ensure the timeliness and accuracy of the knowledge base.
[0094] Detailed description of the process
[0095] User input requirements: Users can ask questions or provide instructions through voice or text input to clearly define the interaction goal. Users can click the "Voice Input" button in the interface to begin speaking. The system supports multiple languages and dialects. Users can also enter specific instructions or questions in the text box, such as "What's the weather like in Beijing tomorrow?" or "Turn on the living room lights." After the user completes their input, they click the "Send" button, and the system automatically proceeds to the next step in the processing flow.
[0096] Speech preprocessing: This cleans, extracts features, and understands the semantics of the input speech, laying the foundation for subsequent processing. The system first performs noise reduction on the uploaded speech, removing any background noise to improve speech quality. Next, a volume adjustment algorithm optimizes speech intelligibility. Subsequently, the system utilizes advanced speech processing techniques to extract key features from the speech signal, such as MFCC, intonation, and speech rate. Furthermore, the system performs preliminary semantic understanding of the speech, identifying the intended category and key information of the voice command, providing basic data support for subsequent in-depth analysis.
[0097] Large-scale model recognition and analysis: A large deep learning model is used to recognize and analyze speech, extracting key information and implicit semantics. The system inputs preprocessed speech into a deep learning model. This model, based on recurrent neural networks (RNNs), Transformer architectures, and other efficient and precise network architectures, can efficiently convert speech signals into text and analyze user intent and contextual information. The model can not only recognize specific commands in speech (such as "turn on the lights"), but also manage multiple rounds of conversations and understand the user's complex needs. Through these in-depth analyses, the system can generate high-quality answers or operational instructions.
[0098] Result Optimization: Optimize the recognition and analysis results, filter important information, and adjust the output format. The system further optimizes the preliminary results generated by the large model to improve the accuracy and usability of the results. First, the system performs semantic optimization on the generated answers or operation instructions to ensure coherent content and clear logic. Then, the system filters out the most important information according to the user's needs, removes redundant content, and highlights the key results. Finally, the system formats and adjusts the results according to the output format specified by the user (such as voice answer, text display, operation feedback, etc.) to meet the diverse needs of users.
[0099] Knowledge Base Assistance: Provide domain knowledge support through the knowledge base to ensure the accuracy and professionalism of the results. The knowledge base module stores rich domain knowledge and data, covering multiple fields such as intelligent customer service, smart home, and in-vehicle voice interaction. The system calls relevant professional knowledge and templates from the knowledge base according to the user's input requirements to provide background information support for the large model. For example, when processing smart home instructions, the system will call the operation guides and related terms of household appliances to ensure that the generated answers are highly professional and accurate. In addition, the knowledge base will be updated in real time according to the user's feedback and the system's operation status to maintain the timeliness and integrity of the knowledge.
[0100] User Interaction and Adjustment: The user adjusts the results according to the needs, and the system responds in real time and updates the results. The user can view the answers or operation feedback generated by the system through the interactive interface and make adjustments as needed. For example, the user can modify the details of the instructions, adjust the tone of the answer, or select a different output format. The system will respond to the user's operations in real time, quickly update the results and redisplay them to ensure that the user can obtain the output that best meets the needs. This interaction mechanism not only improves the user experience but also enhances the flexibility and usability of the system.
[0101] Result Export: The user exports the final results as files in the specified format to complete the entire process. After the user finishes adjusting the results, they can choose to export the results in multiple common formats, such as text files, voice files, operation logs, etc. The user can select the appropriate export format according to the actual needs, and the system will generate the corresponding file according to the selected format and provide a download link or save path. After the user downloads or saves the file, they can apply the voice interaction results generated by the system to actual work or life to complete the entire process.
[0102] Example 2. On the basis of Example 1, a method for an intelligent voice interaction and analysis system based on a large model is proposed, including the following steps: receiving questions or instructions input by the user through voice or text; preprocessing the input voice, including noise reduction, volume adjustment, feature extraction, and preliminary semantic understanding, extracting key features in the voice signal such as Mel Frequency Cepstral Coefficients (MFCC), intonation, and speech rate, and identifying the intent category and key information of the voice instruction; using a deep learning large model to identify and analyze the preprocessed voice, based on recurrent neural networks (RNN), Transformer architecture, and other efficient and accurate network architectures, converting the voice signal into text, parsing the user's intent and context information, and generating a preliminary answer or operation instruction; optimizing the preliminary result generated by the large model, including semantic optimization to ensure coherent content and clear logic, screening important information according to the user's needs, removing redundant content, and typesetting and formatting the result according to the output format specified by the user; according to the user's input requirements, calling relevant professional knowledge and templates from a knowledge base covering multiple fields such as intelligent customer service, smart home, and in-vehicle voice interaction to provide background information support for the large model and assist in generating the final result.
[0103] The user performs voice input by clicking the "Voice Input" button in the interface or directly enters specific instructions or questions in the text box to clarify the interaction target, and then clicks the "Send" button after completion; cleaning the voice input by the user, removing background noise through a filtering algorithm, automatically detecting and adjusting the voice volume using volume normalization technology to enhance voice recognizability, then extracting key features in the voice signal, and classifying the intent of the voice instruction and extracting key information; inputting the preprocessed voice into a deep learning large model, which performs voice recognition to convert the voice signal into text, performs semantic parsing to accurately understand the user's intent and context information, supports multi-turn dialogue management, and generates a preliminary answer or operation instruction; optimizing the preliminary result, performing in-depth semantic optimization on the answer or operation instruction, screening key information according to the user's needs, removing unnecessary redundant information, and performing typesetting adjustment according to the voice answer, text display, and operation feedback output formats specified by the user; the system calls relevant knowledge from a knowledge base storing multi-domain knowledge bases and professional term libraries according to the user's input requirements to provide professional and accurate support for result generation, and the knowledge base is updated in real time to maintain timeliness.
[0104] Receive the needs input by the user in the form of speech or text in multiple languages or dialects; preprocess the input speech, first perform noise reduction and volume adjustment to improve the speech quality, then extract key features such as MFCC, intonation, and speech rate, and at the same time perform semantic understanding on the speech command to determine the intention category such as query, command, or dialogue and extract key information; use a large deep learning model to recognize and analyze the preprocessed speech, convert the speech into text, parse the user's intention and context information, perform multi-round dialogue management, and generate a preliminary result containing speech-to-text conversion, semantic parsing results, answers, or operation feedback; optimize the content of the preliminary result to ensure that the answer is logically rigorous and well-organized, remove redundant information, screen important information according to the user's needs, and perform result typesetting and format adjustment according to the format specified by the user; provide domain knowledge support with the help of a knowledge base. The knowledge base stores knowledge and predefined templates in multiple fields such as intelligent customer service, smart home, and in-vehicle voice interaction, calls relevant knowledge according to the user's input, and the knowledge base is updated in real time according to the user's feedback and system operation conditions.
[0105] The user inputs questions or commands to the system in the form of speech or text, and the system supports multiple languages and dialects; perform a series of preprocessing operations on the input speech, including noise reduction, volume adjustment, feature extraction, and semantic understanding, to obtain the key features of the speech and the command intention and key information; input the preprocessed speech into a large deep learning model based on multiple advanced network architectures, and the model performs speech recognition and semantic parsing, understands the user's intention and context, generates a preliminary answer or operation instruction, and supports multi-round dialogue; perform optimization processing on the preliminary result, perform semantic optimization to make the content coherent and clear, screen key information according to the user's needs, remove redundancy, and perform fine typesetting of the result according to the output format specified by the user; use the knowledge base to provide assistance for result generation. The knowledge base covers knowledge and templates in multiple fields, calls relevant knowledge according to the user's input, and the knowledge base is updated in real time to ensure the accuracy and professionalism of the result. At the same time, it supports the user to create and save custom templates according to the needs.
[0106] Receive the requirements input by the user through voice or text, supporting multiple languages and dialects; preprocess the input voice, including noise reduction, volume adjustment, extraction of MFCC, key features of intonation and speech rate, as well as preliminary semantic understanding to determine the intention category and key information of the voice command; use a large deep learning model to recognize and analyze the preprocessed voice, convert the voice into text, parse the user's intention and context information, conduct multi-round dialogue management, and generate preliminary answers or operation instructions; optimize the preliminary results, perform content semantic optimization, screen important information according to the user's requirements, remove redundant content, and adjust the layout according to the voice answer and text display operation feedback formats specified by the user; provide domain knowledge support with the help of a knowledge base. The knowledge base stores multi-domain knowledge and predefined templates, calls relevant knowledge according to the user input, and the knowledge base is updated in real time according to the user feedback and system operation conditions. Moreover, the system has a model fine-tuning function, which can perform targeted fine-tuning on the large model according to the domain requirements and feedback of the user input to improve the model adaptability and accuracy.
[0107] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent voice interaction and analysis system based on a large model, characterized in that: including: Multi-modal input module: supports hybrid interaction of voice input and text input, is compatible with multiple languages and dialects, and allows users to supplement and correct instructions via voice or text; Real-time feedback mechanism: starts recognition and analysis immediately after receiving input, dynamically displays the processing progress and preliminary results, and supports users to adjust requirements in real time; Multi-module collaborative architecture: integrates voice preprocessing, large model core, result optimization, and knowledge base modules to achieve end-to-end processing from voice signals to executable instructions; Result export function: supports exporting speech-to-text, semantic parsing results, and operation feedback as text, audio, or structured data files; Interactive adjustment interface: provides a graphical operation interface for users to adjust instruction details, optimize answer logic, and select output formats.
2. An intelligent voice interaction and analysis system based on a large model according to claim 1, characterized in that: The voice preprocessing module further includes: Signal cleaning unit: uses an adaptive filtering algorithm to eliminate background noise and unifies the sound pressure level of audio segments through volume normalization technology; Feature extraction unit: extracts feature parameters such as Mel-frequency cepstral coefficients MFCC, fundamental frequency period, and speech rate variation coefficient to construct a high-dimensional representation of the voice signal; Lightweight semantic understanding engine: conducts coarse-grained classification of voice instructions based on a pre-trained language model: query / instruction / dialogue, and extracts key entities such as time, location, and object.
3. The intelligent voice interaction and analysis system based on a large model according to claim 2, characterized in that: The large model core module contains the following innovation points: Multi-granularity speech recognition model: uses a Transformer architecture to fuse the acoustic model and the language model, and supports the recognition of dialect variants of more than 8 languages including Mandarin, Cantonese, and English; Context-aware parser: realizes multi-turn dialogue context tracking through the attention mechanism and automatically fills in omitted information; Domain adaptive fine-tuning framework: supports loading knowledge distillation parameters in professional fields such as medicine and finance, and adjusts the model weights in real time based on user feedback.
4. The intelligent voice interaction and analysis system based on a large model according to claim 3, characterized in that: The result optimization module has the following functions: Semantic refinement algorithm: uses a BERT model to eliminate redundancy and reorganize the logic of the generated text to ensure that the answer conforms to the expression specification of "presenting the conclusion first - hierarchical argumentation"; Multi-modal output adaptation layer: automatically converts the format according to the output scenario, such as generating a technical document in Markdown format, a device control instruction in JSON structure, or an emotional speech synthesis waveform; Confidence screening mechanism: sorts the parsing results by confidence, only presents information above the threshold, and marks the possible error types of low-confidence content.
5. An intelligent voice interaction and analysis system based on a large model according to claim 4, characterized in that: The knowledge base module achieves the following technical effects: Hybrid knowledge graph: constructs a three-layer knowledge architecture including a general encyclopedia, a domain term library, and a scenario template, where the smart home template contains a triple relationship of "device-action-environment"; Incremental update mechanism: automatically discovers new knowledge by monitoring user dialogue logs, and uses a graph neural network GNN to achieve the dynamic evolution of knowledge embedding vectors; Template-based response generation: pre-sets more than 50 scenario response templates and supports users to customize parameter mapping rules.
6. A method for an intelligent voice interaction and analysis system based on a large model according to claim 5, characterized in that: including the following steps: Receive questions or instructions input by the user via voice or text; Preprocess the input speech, including noise reduction, volume adjustment, feature extraction, and preliminary semantic understanding. Extract key features in the speech signal such as Mel Frequency Cepstral Coefficients (MFCC), intonation, and speech rate, and identify the intent category and key information of the speech command. Use a large deep learning model to recognize and analyze the preprocessed speech. Based on recurrent neural network (RNN), Transformer architecture, and other efficient and accurate network architectures, convert the speech signal into text, parse the user's intent and context information, and generate a preliminary response or operation instruction. Optimize the preliminary results generated by the large model, including semantic optimization to ensure content coherence and logical clarity, screen important information according to user needs, remove redundant content, and format and adjust the results according to the output format specified by the user. According to the user's input requirements, call relevant professional knowledge and templates from a knowledge base covering multiple fields such as intelligent customer service, smart home, and in-vehicle voice interaction to provide background information support for the large model and assist in generating the final result.
7. A method according to claim 6, characterized in that: Include the following steps: The user performs voice input by clicking the "Voice Input" button in the interface or directly enters specific instructions or questions in the text box, clarifies the interaction goal, and then clicks the "Send" button after completion. Clean the user's input speech. Remove background noise through a filtering algorithm, automatically detect and adjust the speech volume using volume normalization technology to enhance speech recognizability. Then extract the key features in the speech signal, and classify the intent of the speech command and extract key information. Input the preprocessed speech into a large deep learning model. The model performs speech recognition to convert the speech signal into text, performs semantic parsing to accurately understand the user's intent and context information, supports multi-turn dialogue management, and generates a preliminary response or operation instruction. Optimize the preliminary results. Perform in-depth semantic optimization on the response or operation instruction, screen key information according to user needs, remove unnecessary redundant information, and perform typesetting adjustment according to the output formats of voice response, text display, and operation feedback specified by the user. The system calls relevant knowledge from a knowledge base storing knowledge bases and professional term libraries in multiple fields to provide support for professionalism and accuracy in result generation. The knowledge base is updated in real time to maintain timeliness.
8. A method according to claim 6, wherein: Include the following steps: Receive the user's requirements input in various languages or dialects through voice or text. Preprocess the input speech. First, perform noise reduction and volume adjustment to improve speech quality, then extract key features such as MFCC, intonation, and speech rate. At the same time, perform semantic understanding on the speech command to determine the intent category such as query, instruction, or dialogue and extract key information. Use a large deep learning model to recognize and analyze the preprocessed speech, convert the speech into text, parse the user's intent and context information, perform multi-turn dialogue management, and generate a preliminary result including speech-to-text conversion, semantic parsing results, response, or operation feedback. Optimize the content of the preliminary results to ensure rigorous logic and clear organization in the answers, remove redundant information, screen important information according to the user's needs, and perform result typesetting and format adjustment according to the user-specified format; Leverage the knowledge base to provide domain knowledge support. The knowledge base stores knowledge and predefined templates in multiple fields such as intelligent customer service, smart home, and in-vehicle voice interaction, calls relevant knowledge according to user input, and the knowledge base is updated in real time based on user feedback and system operation status.
9. A method according to claim 6, wherein: It includes the following steps: The user inputs questions or instructions to the system in the form of voice or text, and the system supports multiple languages and dialects; Perform a series of preprocessing operations on the input voice, including noise reduction, volume adjustment, feature extraction, and semantic understanding, to obtain the key features of the voice, the intention of the instruction, and key information; Input the preprocessed voice into a deep learning large model based on multiple advanced network architectures. The model performs speech recognition and semantic parsing, understands the user's intention and context, generates a preliminary answer or operation instruction, and supports multi-round conversations; Optimize the preliminary results, perform semantic optimization to make the content coherent and clear, screen key information according to the user's needs, remove redundancy, and perform fine typesetting on the results according to the user-specified output format; Utilize the knowledge base to provide assistance for result generation. The knowledge base covers knowledge and templates in multiple fields, calls relevant knowledge according to user input, and the knowledge base is updated in real time to ensure the accuracy and professionalism of the results. At the same time, the system supports users to create and save custom templates according to their needs.
10. A method according to claim 6, characterized in that: It includes the following steps: Receive the user's needs input in the form of voice or text, and support multiple languages and dialects; Preprocess the input voice, including noise reduction, volume adjustment, extraction of MFCC, key features of intonation and speech rate, and preliminary semantic understanding, to determine the intention category and key information of the voice instruction; Use a deep learning large model to recognize and analyze the preprocessed voice, convert the voice into text, parse the user's intention and context information, manage multi-round conversations, and generate a preliminary answer or operation instruction; Optimize the preliminary results, perform content semantic optimization, screen important information according to the user's needs, remove redundant content, and perform typesetting adjustment according to the user-specified voice answer, text display operation feedback format; Leverage the knowledge base to provide domain knowledge support. The knowledge base stores knowledge and predefined templates in multiple fields, calls relevant knowledge according to user input, the knowledge base is updated in real time based on user feedback and system operation status, and the system has a model fine-tuning function, which can perform targeted fine-tuning on the large model according to the domain needs and feedback of user input to improve the adaptability and accuracy of the model.
Citation Information
Cited By
Knowledge extraction and mining method for smart home multi-mode dialogue system
CN120850235A
Voice control method and voice control system
CN121171225A
AI intelligent sound equipment signal processing method and system, sound equipment device and storage medium
CN121789667A