Voice and video seat auxiliary prompting method based on large model

The method addresses the limitations of static rule-based systems by integrating multi-modal data processing and context understanding to enhance the responsiveness and personalization of seat assistance, improving user experience and service efficiency.

CN120319230AActive Publication Date: 2025-07-15北京微呼科技有限公司

Patent Information

Application Number
CN202510822602.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-15
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

In the prior art, the agent assist system lacks the joint modeling and real-time context perception ability of multimodal interaction characteristics, which makes it difficult to take into account the quality of prompts and timeliness, affecting customer service efficiency and user experience.

Method used

The voice and video seat assisted prompt method based on the big model is adopted, and multimodal data processing module, context understanding module and personalized service generation module are combined with real-time optimization mechanism to realize multimodal feature fusion, context analysis and personalized prompt content generation.

Benefits of technology

It improves the efficiency and accuracy of auxiliary prompts in voice and video seat scenarios, enhances user experience, supports continuous improvement of the system and personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319230A_ABST
    Figure CN120319230A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice and video seat assistance, in particular to a voice and video seat assistance prompting method based on a large model. The method comprises the steps of collecting user session data to construct a multi-modal feature matrix, generating and analyzing a context vector, generating personalized prompt content and optimizing an output priority, transmitting the prompt content to a seat end, and feeding back an update model. According to the application, the context analysis capability in a complex scene is improved through the multi-modal data processing and context understanding module, the prompt accuracy and flexibility are improved in combination with the personalized service generation and real-time optimization module, and the continuous improvement of the system is realized by using the feedback optimization module; therefore, the response speed, the prompt accuracy and the personalized service level in voice and video seat scenes are remarkably improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and customer service, and specifically relates to a speech and video seat assistance prompting method based on a large model. Background Art

[0002] Intelligent interaction technology refers to using artificial intelligence, natural language processing, machine learning, and other related technologies to enhance the interaction method between humans and computer systems. When intelligent interaction technology is applied to the speech and video seat assistance prompting method, it can improve customer service efficiency and optimize the user experience.

[0003] The existing technology has the following deficiencies:

[0004] Currently, most seat assistance systems still rely on static rule libraries or template-driven prompting mechanisms, lacking the ability to jointly model multi-modal interaction features (such as speech intonation, facial expressions, text semantics, etc.) and real-time context awareness, and it is difficult to dynamically adapt to complex and changing user intentions. It is difficult to balance the prompting quality and timeliness, which affects customer service efficiency and user experience. Therefore, a speech and video seat assistance prompting method based on a large model is proposed.

[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, so it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] In order to overcome the above defects of the existing technology, the present invention provides a speech and video seat assistance prompting method based on a large model. By introducing a multi-modal data processing module, a context understanding module, and a personalized service generation module, combined with a real-time optimization mechanism, it solves the problems of slow response speed, insufficient prompting accuracy, and limited personalized service ability in the existing technology in complex scenarios. The present invention can achieve more efficient assistance prompting in the speech and video seat scenario and improve the user experience.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A speech and video seat assistance prompting method based on a large model, comprising the following steps:

[0009] Comprising the following steps:

[0010] Step S1: Collect user session data and construct a multi-modal feature matrix, and set a time window to capture dynamic interaction information;

[0011] Step S2: Generate a context vector using the multi-modal feature matrix, and perform in-depth parsing of the current session in combination with the context understanding module;

[0012] Step S3: Generate personalized prompt content based on the context vector, and adjust the output priority through the real-time optimization module;

[0013] Step S4: Transmit the generated prompt content to the agent terminal, and update the prompt generation model according to the feedback information.

[0014] In a preferred embodiment, in step S1, the user session data includes voice signals, video signals, and text inputs;

[0015] The voice signal is converted into a spectrogram feature map through an acoustic feature extraction module;

[0016] The video signal extracts the key point positions and emotion features through a facial expression recognition module;

[0017] The text input is converted into a word vector sequence through a natural language processing module;

[0018] The voice signal, video signal, and text input are respectively stored in independent feature buffer areas, and merged into a multi-modal feature matrix through a feature fusion algorithm. The feature fusion algorithm uses the weighted average method, and the weights are determined by the feature contribution degrees in the historical data.

[0019] In a preferred embodiment, in step S1, the setting of the time window is implemented through a sliding window mechanism. The length of the sliding window is a fixed value, and the step size is dynamically adjusted according to the real-time load;

[0020] When the system detects high-concurrency requests, shorten the step size of the sliding window to reduce the consumption of computing resources;

[0021] When the system load is low, extend the step size of the sliding window to improve the feature capture accuracy. The data within the sliding window is converted into a time-dimensional feature vector through a time series encoder.

[0022] In a preferred embodiment, in step S2, the multi-modal feature matrix is input into the context understanding module. The context understanding module includes a three-layer neural network structure:

[0023] The input layer receives the multi-modal feature matrix, normalizes it, and then inputs it into the attention layer. The attention layer calculates the importance weights of each feature through a self-attention mechanism;

[0024] The formula is ;

[0025] where, 、 and are respectively the query matrix, key matrix, and value matrix of the input features, is the dimension of the key matrix, and the output layer multiplies the attention weights by the input features to generate a context vector.

[0026] In a preferred embodiment, in step S2, after the context vector is generated, the current conversation is deeply parsed by a context parsing algorithm, and the context parsing algorithm adopts a hierarchical analysis method;

[0027] First, the context vector is subject classified to determine the topic category of the current conversation. Secondly, the topic category is refined by combining the context information to generate fine-grained context labels. The fine-grained context labels are stored in the context label library and compared with the historical conversation data through a context matching algorithm.

[0028] In a preferred embodiment, in step S3, the generation of personalized prompt content is implemented by a prompt generation module. The prompt generation module includes two parts: a template generator and a dynamic generator. The template generator generates basic prompt content based on predefined templates in the context label library. The dynamic generator combines the context vector and the user profile to generate customized prompt content. The user profile is generated by a historical interaction data analysis module and includes the user's preferences, behavior patterns, and emotional tendencies.

[0029] In a preferred embodiment, in step S3, the real-time optimization module adjusts the output order of the prompt content through a priority scheduling algorithm. The priority scheduling algorithm comprehensively considers the importance and urgency of the prompt content. The importance is determined by the weight of the context vector, and the urgency is determined by the interaction frequency within the time window. The priority calculation formula is P = α·W + β·F;

[0030] where P is the priority of the prompt content, W is the weight of the context vector, F is the interaction frequency, and α and β are weight coefficients.

[0031] In a preferred embodiment, in step S4, the generated prompt content is transmitted to the agent terminal through the agent terminal interface. The agent terminal interface supports the transmission of multiple data formats, including text, images, and voices. The display method of the prompt content is dynamically adjusted according to the type of the agent terminal device. The system obtains the operation records and user evaluations of the agent terminal through a feedback collection module to update the prompt generation model.

[0032] In a preferred embodiment, in step S4, the feedback collection module obtains feedback information through a multi-channel data collection mechanism, including the operation logs of the agent terminal, the satisfaction scores of users, and the emotional characteristics in voice and video signals;

[0033] The feedback information is stored in the feedback database and the prompt generation model is updated through an incremental learning algorithm. The incremental learning algorithm adopts the mini-batch gradient descent method, and only the latest feedback data is used for each update.

[0034] A voice and video agent assistance prompt system based on a large model includes a data acquisition module, a context understanding module, a prompt generation module, and a feedback optimization module;

[0035] The data acquisition module is used to collect user session data and construct a multimodal feature matrix, and transmit the multimodal feature matrix to the context understanding module;

[0036] The context understanding module is used to generate a context vector and deeply analyze the current session, and transmit the analysis result to the prompt generation module;

[0037] The prompt generation module is used to generate personalized prompt content and adjust the output priority through a real-time optimization module, and finally transmit the prompt content to the agent side;

[0038] The feedback optimization module is used to update the prompt generation model by collecting feedback information to form a closed-loop optimization mechanism.

[0039] Technical effects and advantages of the present invention:

[0040] By introducing a multimodal data processing module and a context understanding module, the present invention solves the problem of insufficient context understanding ability of the prior art in complex scenarios; through a personalized service generation module and a real-time optimization module, the flexibility and accuracy of the prompt content are improved; through the feedback optimization module, continuous improvement of the system is achieved, meeting the actual needs of the voice and video agent scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic flowchart of a method for voice and video agent assistance prompt based on a large model of the present invention.

[0042] Figure 2 It is a module structure diagram of a voice and video agent assistance prompt system based on a large model of the present invention.

[0043] Figure 3 It is a schematic diagram of a sliding window mechanism in a method for voice and video agent assistance prompt based on a large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] The present invention provides a method and system for voice and video agent assistance prompt based on a large model, and its specific implementation is as Figures 1 to 3As shown below; the specific implementation of the present invention will be described in detail in conjunction with the accompanying drawings to ensure that those skilled in the art can implement this technology based on the content of the specification.

[0046] As Figure 1 shown, the overall process of the present invention includes four main links: data collection, context understanding, prompt generation, and feedback optimization, which are collaboratively completed by the data collection module, context understanding module, prompt generation module, and feedback optimization module respectively.

[0047] These modules are interconnected in the form of a data flow to form a closed-loop system, thereby realizing the complete process from the collection of user session data to the generation of prompt content and model optimization.

[0048] In step S1, the data collection module is responsible for collecting user session data and constructing a multi-modal feature matrix.

[0049] User session data includes three types: voice signals, video signals, and text inputs. Voice signals are converted into spectrogram features through an acoustic feature extraction module, video signals have key point positions and emotion features extracted through a facial expression recognition module, and text inputs are converted into word vector sequences through a natural language processing module.

[0050] The above three types of features are stored in independent feature buffer areas respectively, and are combined into a multi-modal feature matrix through a feature fusion algorithm.

[0051] The feature fusion algorithm adopts the weighted average method, and the weights are determined by the feature contribution degrees in historical data. The setting of the time window is realized through a sliding window mechanism;

[0052] As Figure 3 shown, the length of the sliding window is a fixed value, and the step size can be dynamically adjusted according to the real-time load. When the system detects high-concurrency requests, the step size of the sliding window is shortened to reduce the consumption of computing resources;

[0053] When the system load is low, the step size of the sliding window is extended to improve the feature capture accuracy.

[0054] The data within the sliding window is converted into time-dimensional feature vectors through a time series encoder for the generation of subsequent context vectors.

[0055] The data collection module transfers the generated multi-modal feature matrix to the context understanding module.

[0056] Realize collaborative modeling of multimodal data: Collect three types of conversation data, namely user voice signals, video signals, and text inputs, and convert them into spectral feature maps, emotion key point features, and word vector sequences respectively through modules such as acoustic feature extraction, facial expression recognition, and natural language processing, so as to construct a multimodal feature set with speech, visual, and semantic dimensions, and enhance the input data dimension and information density for downstream context understanding.

[0057] Enhance the robustness and dynamic adaptability of feature fusion: Adopt a weighted average fusion strategy, and dynamically adjust the weight coefficients based on the feature contribution degree statistically calculated from historical data, so that the multimodal feature fusion process has a semantic association strengthening mechanism, can achieve the optimal combination of feature information in different user scenarios, and improve the discriminant performance of the fused feature matrix.

[0058] Improve the system's real-time response and resource regulation capabilities: Integrate a sliding window mechanism to batch process multimodal data within a time segment. The window length is fixed and the step size is adjustable. The system automatically shortens the step size to reduce resource consumption under high load, and extends the step size to increase data granularity and analysis accuracy under low load, taking into account both system stability and feature capture integrity.

[0059] Support time series modeling capabilities: Perform sequence modeling on the data within the sliding window through a time series encoder to generate time feature vectors with time dimension dependencies, providing stable and temporally consistent data inputs for subsequent context vector generation and user intention recognition, and enhancing the continuity and context consistency of context awareness.

[0060] In step S2, the context understanding module receives the multimodal feature matrix from the data collection module and processes it through a three-layer neural network structure.

[0061] The input layer receives the multimodal feature matrix, normalizes it and then passes it to the attention layer; the attention layer calculates the importance weights of each feature through the self-attention mechanism;

[0062] The formula is ;

[0063] Among them, , and are the query matrix, key matrix, and value matrix of the input features respectively, is the dimension of the key matrix, and the output layer multiplies the attention weights by the input features to generate a context vector, which is used to represent the global context information of the current conversation.

[0064] The generated context vector is then deeply analyzed by a context parsing algorithm. The context parsing algorithm adopts a hierarchical analysis method. First, it classifies the context vector by theme to determine the theme category of the current conversation. Second, it refines the theme category by combining the context information to generate fine-grained context tags.

[0065] The fine-grained context tags are stored in the context tag library and compared with the historical conversation data through a context matching algorithm to improve the accuracy of context parsing.

[0066] The context understanding module transmits the generated context vector and fine-grained context tags to the prompt generation module.

[0067] Enhance the semantic extraction ability of multi-modal features: Use a three-layer neural network structure to process the multi-modal feature matrix from the data acquisition module. The input layer implements a normalization operation to ensure the consistency of features in the numerical scale. The middle layer introduces a self-attention mechanism to calculate the importance weights of each modal feature in the current conversation context, effectively highlighting the feature dimensions that make key contributions to the global semantics, and improving the accuracy and robustness of the overall semantic extraction.

[0068] Introduce a self-attention mechanism to improve the context modeling accuracy: By constructing query matrices, key matrices, and value matrices, and performing feature weighting according to the following attention mechanism formula.

[0069] Implement hierarchical parsing of context vectors: The context vector generated by the output layer not only contains multi-modal fusion information but also has the global perception ability of the context dimension. This vector is hierarchically processed by the context parsing algorithm. First, it realizes the context classification at the theme level, and then combines the historical context to be refined into more recognizable fine-grained context tags, effectively improving the system's semantic understanding ability in multi-turn and multi-topic alternating scenarios.

[0070] Enhance the dynamic adaptation and evolution ability of context tags: By constructing a context tag library and combining the context matching algorithm to conduct a comparative analysis of historical conversation data, realize the automatic comparison and dynamic fusion of the current context with the existing tag system, thereby improving the intelligent level of tag update and adaptation, and further optimizing the response logic of the subsequent prompt generation module.

[0071] Improve the intelligent interaction quality of the overall system: This step not only enhances the comprehensive interpretation ability of the context understanding module for multi-modal data but also provides the prompt generation module with deep context information with clear structure, precise semantics, and consistent context, providing a solid data foundation for realizing personalized and strongly semantic-consistent system responses.

[0072] In step S3, the prompt generation module generates personalized prompt content according to the context vector and fine-grained context tags. The prompt generation module includes two parts: a template generator and a dynamic generator;

[0073] The template generator generates basic prompt content based on predefined templates in the context tag library, while the dynamic generator generates customized prompt content by combining context vectors and user profiles.

[0074] The user profile is generated by the historical interaction data analysis module and contains information such as the user's preferences, behavior patterns, and emotional tendencies.

[0075] The generated prompt content is adjusted for output priority by the real-time optimization module. The real-time optimization module comprehensively considers the importance and urgency of the prompt content through a priority scheduling algorithm. The importance is determined by the weight of the context vector, and the urgency is determined by the interaction frequency within the time window.

[0076] The priority calculation formula is P = α·W + β·F;

[0077] where P is the priority of the prompt content, W is the weight of the context vector, F is the interaction frequency, and α and β are weight coefficients that can be determined through experimental tuning.

[0078] The prompt generation module passes the generated prompt content to the feedback optimization module.

[0079] Enhancing the personalization and pertinence of prompt content: The prompt generation module integrates the template generator and the dynamic generator. The former generates efficient and standardized basic prompt content based on structured corpus templates in the context tag library, while the latter customizes content by combining real-time context vectors and user profile information, fully considering the user's historical preferences, behavior patterns, and emotional tendencies, to achieve personalized prompt generation that links context awareness and user intent, significantly improving the adaptability of system prompts and user satisfaction.

[0080] Achieving context-enhanced output driven by user profiles: Constructing a multi-dimensional user profile through the historical interaction data analysis module and performing correlation analysis with the current context vector to enhance the adaptability of prompt content to the current user state, supporting cross-conversation context maintenance and content continuity modeling, and being applicable to complex application scenarios such as multi-turn Q&A and multi-modal interaction.

[0081] Improving the real-time performance and response efficiency of prompt scheduling: Introducing a real-time optimization module to dynamically sort the generated prompt content and adopting a priority scheduling formula.

[0082] Balancing the stability of structural templates and the flexibility of semantic generation: The dual-generator architecture effectively combines the generality of structured language templates and the flexibility of neural generation models, ensuring both the consistency and language quality of prompt content while also having diversity and adaptability, solving the problem of severe templatization and lack of variation in prompt content in traditional systems.

[0083] Enhance the interactive guidance and behavior incentive capabilities of the system: The generated prompt content is not only output as passive response items, but can also be actively used to guide user behavior, stimulate user participation, and correct user paths, thus constructing a more proactive and guiding interactive loop in scenarios such as intelligent assistants, medical recommendations, and educational interactions.

[0084] In step S4, the feedback optimization module transmits the prompt content to the agent terminal through the agent terminal interface, and the agent terminal interface supports the transmission of multiple data formats, including text, images, and voices.

[0085] The display method of the prompt content is dynamically adjusted according to the type of the agent terminal device. For example, it is displayed in the form of a pop-up window on a desktop device and in the form of a notification bar on a mobile device.

[0086] At the same time, the system obtains the operation records and user evaluations of the agent terminal through the feedback collection module for updating the prompt generation model.

[0087] Achieve multi-terminal adaptive distribution of prompt content: The feedback optimization module transmits the prompt content to different types of agent terminals through the agent terminal interface, and this interface supports the dynamic transmission of multiple data formats such as text, images, and voices, ensuring the compatibility and display consistency of the prompt content on multiple terminals such as desktop terminals, mobile terminals, and wearable devices, and significantly improving the flexibility of system deployment and cross-platform operation capabilities.

[0088] Enhance the context awareness ability of prompt display: The presentation form of the prompt content can be dynamically adjusted according to the type of the agent terminal device. For example, a desktop device uses a pop-up prompt to enhance attention capture, and a mobile device uses a notification bar prompt to reduce interference, so as to match the display method with the user's usage habits in different interaction scenarios, and improve the acceptance rate and response efficiency of prompt information.

[0089] Build a closed-loop feedback mechanism for prompt generation: Real-time collect the operation records and user evaluations of the agent terminal through the feedback collection module, including indicators such as click behavior, ignore frequency, and satisfaction score, so as to establish a response mapping relationship between prompt output and user behavior, feed back the prompt generation model and continuously optimize the accuracy and applicability of the prompt content, and promote the system to evolve from static recommendation to dynamic adaptation.

[0090] Improve the self-learning and dynamic evolution capabilities of the system: The feedback information, as an important training sample of the prompt generation model, is embedded in the model training process through a periodic update mechanism, enabling the system to continuously adjust the content generation strategy based on real interaction data, realizing the evolution of the model from rule-driven to data-driven, and improving the intelligence level and robustness of the system under long-term operation.

[0091] Support system performance evaluation and policy iteration: The structured collection of operation records and user evaluations helps to construct an evaluation index system for the system prompt effect, such as prompt response rate, prompt hit rate, user satisfaction curve, etc., providing a quantitative basis for prompt policy optimization, interface design improvement and model parameter fine-tuning, and supporting the debuggability and portability of the system in different business fields.

[0092] The feedback collection module obtains feedback information through a multi-channel data collection mechanism, including operation logs at the agent end, user satisfaction scores, and emotional features in voice and video signals.

[0093] The feedback information is stored in the feedback database and the prompt generation model is updated through an incremental learning algorithm. The incremental learning algorithm adopts the mini-batch gradient descent method and only uses the latest feedback data for each update to reduce the computational overhead.

[0094] The feedback optimization module feeds the updated prompt generation model back to the prompt generation module to form a closed-loop optimization mechanism.

[0095] The data collection module, context understanding module, prompt generation module, and feedback optimization module are tightly connected in the form of a data stream.

[0096] The data collection module transfers the multi-modal feature matrix to the context understanding module, the context understanding module transfers the generated context vector and fine-grained context labels to the prompt generation module, the prompt generation module transfers the generated prompt content to the feedback optimization module, and the feedback optimization module feeds the updated prompt generation model back to the prompt generation module.

[0097] This collaborative relationship between modules ensures the efficient operation and continuous optimization of the system.

[0098] In practical applications, the present invention can be widely applied to voice and video agent scenarios, such as customer service, technical support, and telemedicine.

[0099] Taking customer service as an example, when a user interacts with an agent through voice or video, the data collection module real-time collects the user's voice signal, facial expression, and text input, and generates a multi-modal feature matrix through a feature fusion algorithm.

[0100] The context understanding module deeply analyzes the multi-modal feature matrix to generate a context vector and fine-grained context labels.

[0101] The prompt generation module generates personalized prompt content according to the context vector and user profile, such as recommending solutions or providing operation guidance, and adjusts the output priority through a real-time optimization module.

[0102] The feedback optimization module continuously optimizes the prompt generation model by collecting the operation records and user evaluations at the agent end, thereby improving the system's response speed and prompt accuracy.

[0103] By introducing a multi-modal data processing module and a context understanding module, the present invention solves the problem of insufficient context understanding ability in the prior art in complex scenarios;

[0104] Through the personalized service generation module and the real-time optimization module, the flexibility and accuracy of the prompt content are improved;

[0105] Through the feedback optimization module, continuous improvement of the system is achieved, meeting the actual needs of the voice and video agent scenarios.

[0106] To enable the relevant personnel in the technical field to fully understand and implement the present invention better, the following further supplements and explains the specific implementation principle of the present invention in combination with a specific application scenario.

[0107] In the telemedicine scenario, when a patient interacts with a doctor through a video agent, the system can assist the doctor in real time to provide more accurate diagnostic suggestions.

[0108] First, the data acquisition module respectively collects the patient's voice signal and video signal through a microphone and a camera, and synchronously receives the text description information input by the patient.

[0109] The voice signal is processed by the acoustic feature extraction module and converted into a spectrogram feature map, while the video signal extracts the facial key point positions and emotional features of the patient through the facial expression recognition module, such as pain expressions or anxiety states.

[0110] The text input is converted into a word vector sequence through the natural language processing module to capture semantic information. The above three types of feature data are respectively stored in independent feature buffer areas and merged into a multi-modal feature matrix through a feature fusion algorithm.

[0111] The feature fusion algorithm adopts the weighted average method, and the weights are determined by the feature contribution degrees in the historical data, so as to ensure that the multi-modal feature matrix can comprehensively reflect the patient's current state.

[0112] Subsequently, the context understanding module receives the multi-modal feature matrix from the data acquisition module and processes it through a three-layer neural network structure.

[0113] The input layer standardizes the multi-modal feature matrix and then passes it to the attention layer, and the attention layer calculates the importance weights of each feature through the self-attention mechanism.

[0114] For example, when the patient describes symptoms, the tone change in the voice signal may be more representative than the text input, so its weight will be dynamically increased.

[0115] The output layer multiplies the attention weights with the input features to generate a context vector, which is used to represent the global context information of the current conversation.

[0116] The generated context vector is further deeply parsed through a context parsing algorithm. First, the context vector is classified by topic, for example, to determine whether the topic of the current conversation is "acute diseases" or "chronic disease management".

[0117] Then, the topic category is refined by combining the context information to generate fine-grained context labels, such as "acute abdominal pain" or "hypertension follow-up".

[0118] These fine-grained context labels are stored in the context label library and compared with the historical conversation data through a context matching algorithm to improve the accuracy of context parsing.

[0119] The prompt generation module generates personalized prompt content based on the context vector and the fine-grained context labels.

[0120] The template generator generates basic prompt content based on predefined templates in the context label library, such as common diagnostic suggestions for "acute abdominal pain".

[0121] The dynamic generator combines the context vector and the patient profile to generate customized prompt content.

[0122] The patient profile is generated by the historical interaction data analysis module and contains information such as the patient's past medical history, medication records, and consultation preferences.

[0123] For example, for a patient with a history of gastric ulcer, the system will preferentially recommend prompt content related to gastric examinations.

[0124] The generated prompt content is adjusted for output priority by the real-time optimization module.

[0125] The real-time optimization module comprehensively considers the importance and urgency of the prompt content through a priority scheduling algorithm.

[0126] For example, when the symptoms described by the patient are relatively severe and the interaction frequency is high, the priority of the prompt content will be significantly increased.

[0127] The priority calculation formula is P = α·W + β·F;

[0128] where P is the priority of the prompt content, W is the weight of the context vector, F is the interaction frequency, and α and β are weight coefficients that can be determined through experimental tuning.

[0129] The feedback optimization module transmits the prompt content to the doctor's device through the agent interface. On a desktop device, the prompt content is displayed in the form of a pop-up window; on a mobile device, it is presented in the form of a notification bar.

[0130] Meanwhile, the system obtains the operation records of doctors and the satisfaction scores of patients through the feedback collection module.

[0131] For example, whether the doctor has adopted the prompt content of the system, and the patient's satisfaction evaluation of the diagnosis process.

[0132] In addition, the system also analyzes the feedback information of patients through the emotional features in voice and video signals. These feedback information are stored in the feedback database and the prompt generation model is updated through the incremental learning algorithm.

[0133] The incremental learning algorithm adopts the mini-batch gradient descent method and only uses the latest feedback data for each update to reduce the computational overhead.

[0134] For example, when the system finds that the adoption rate of a certain type of prompt content is low, it will automatically adjust the relevant weights to optimize the generation strategy of the prompt content.

[0135] Through the above steps, the present invention realizes an efficient auxiliary prompt function in the telemedicine scenario. The data acquisition module captures the multi-modal data of patients in real time, the context understanding module deeply analyzes the complex context, the prompt generation module generates personalized prompt content, and the feedback optimization module continuously improves the system performance through a closed-loop mechanism.

[0136] This close cooperation between modules ensures the efficient operation of the system, which not only improves the work efficiency of doctors, but also significantly improves the medical experience of patients.

[0137] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0138] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application of the technical solution and the invention constraints. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application.

[0139] In addition, each functional module in the various embodiments of the present application can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0140] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.

[0141] Finally: The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A voice and video seat assistance prompting method based on a large model, characterized in that, It includes the following steps: Step S1: Collect user session data and construct a multi-modal feature matrix, and set a time window to capture dynamic interaction information; Step S2: Generate a context vector using the multi-modal feature matrix, and deeply analyze the current session in combination with the context understanding module; Step S3: Generate personalized prompt content based on the context vector, and adjust the output priority through the real-time optimization module; Step S4: Transmit the generated prompt content to the agent terminal, and update the prompt generation model according to the feedback information.

2. A voice and video agent assistance prompt method based on a large model according to claim 1, characterized in that: In step S1, the user session data includes voice signals, video signals, and text inputs; The voice signal is converted into a spectrogram feature map through an acoustic feature extraction module; The video signal extracts key point positions and emotion features through a facial expression recognition module; The text input is converted into a word vector sequence through a natural language processing module; The voice signal, video signal, and text input are respectively stored in independent feature buffer areas, and merged into a multi-modal feature matrix through a feature fusion algorithm. The feature fusion algorithm uses the weighted average method, and the weights are determined by the feature contribution degrees in the historical data.

3. A voice and video agent assistance prompt method based on a large model according to claim 2, characterized in that: In step S1, the setting of the time window is realized through a sliding window mechanism. The length of the sliding window is a fixed value, and the step size is dynamically adjusted according to the real-time load; When the system detects a high-concurrency request, shorten the step size of the sliding window to reduce the consumption of computing resources; When the system load is low, extend the step size of the sliding window to improve the feature capture accuracy. The data within the sliding window is converted into a time-dimensional feature vector through a time series encoder.

4. A voice and video agent assistance prompt method based on a large model according to claim 3, characterized in that: In step S2, the multi-modal feature matrix is input into the context understanding module. The context understanding module includes a three-layer neural network structure: The input layer receives the multi-modal feature matrix, normalizes it, and then passes it to the attention layer. The attention layer calculates the importance weights of each feature through the self-attention mechanism; The formula is ; Among them, , and are the query matrix, key matrix, and value matrix of the input features respectively. is the dimension of the key matrix, and the output layer generates a context vector by multiplying the attention weights with the input features.

5. A voice and video agent assistance prompt method based on a large model according to claim 4, characterized in that: In step S2, after the context vector is generated, the current session is deeply analyzed through a context parsing algorithm. The context parsing algorithm adopts a hierarchical analysis method; First, the context vector is subject classified to determine the theme category of the current session. Secondly, the theme category is refined in combination with the context information to generate fine-grained context labels. The fine-grained context labels are stored in the context label library and compared with the historical session data through a context matching algorithm.

6. A voice and video agent assistance prompt method based on a large model according to claim 5, characterized in that: In step S3, the generation of personalized prompt content is implemented by a prompt generation module. The prompt generation module includes two parts: a template generator and a dynamic generator. The template generator generates basic prompt content based on predefined templates in the context tag library, and the dynamic generator generates customized prompt content by combining the context vector and the user profile. The user profile is generated by the historical interaction data analysis module and includes the user's preferences, behavior patterns, and emotional tendencies.

7. A speech and video seat assistance prompt method based on a large model according to claim 6, characterized in that: In step S3, the real-time optimization module adjusts the output order of the prompt content through a priority scheduling algorithm. The priority scheduling algorithm comprehensively considers the importance and urgency of the prompt content. The importance is determined by the weight of the context vector, and the urgency is determined by the interaction frequency within the time window. The priority calculation formula is P = α·W + β·F; Where P is the priority of the prompt content, W is the weight of the context vector, F is the interaction frequency, and α and β are weight coefficients.

8. A speech and video seat assistance prompt method based on a large model according to claim 7, characterized in that: In step S4, the generated prompt content is transmitted to the seat end through the seat end interface. The seat end interface supports the transmission of multiple data formats, including text, image, and voice. The display method of the prompt content is dynamically adjusted according to the type of the seat end device. The system obtains the operation records and user evaluations of the seat end through the feedback collection module to update the prompt generation model.

9. A speech and video seat assistance prompt method based on a large model according to claim 8, characterized in that: In step S4, the feedback collection module obtains feedback information through a multi-channel data collection mechanism, including the operation logs of the seat end, the satisfaction scores of users, and the emotional characteristics in the speech and video signals; The feedback information is stored in the feedback database and the prompt generation model is updated through an incremental learning algorithm. The incremental learning algorithm uses the mini-batch gradient descent method, and only the latest feedback data is used for each update.

10. A voice and video seat auxiliary prompt system based on a large model, which is used to implement a voice and video seat auxiliary prompt method according to any one of claims 1-9, and is characterized in that: It includes a data collection module, a context understanding module, a prompt generation module, and a feedback optimization module; The data collection module is used to collect user session data and construct a multi-modal feature matrix, and transmit the multi-modal feature matrix to the context understanding module; The context understanding module is used to generate a context vector and deeply analyze the current session, and transmit the analysis result to the prompt generation module; The prompt generation module is used to generate personalized prompt content and adjust the output priority through the real-time optimization module, and finally transmit the prompt content to the seat end; The feedback optimization module is used to update the prompt generation model by collecting feedback information to form a closed-loop optimization mechanism.

Citation Information

Patent Citations

  • Generative question answering system, method and equipment, storage medium and program product

    CN117909478A

  • Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search

    US20240386015A1

  • System and method for a dialogue response generation system

    WO2021049199A1

  • Text semantic matching method and refrigeration device system

    WO2024188277A1

  • Multi-modal knowledge-based question answering method and system for 5g message

    WO2025060773A1

Cited By

  • Real-time verbal skill decision-making method and system based on multi-mode and session state perception

    CN120950661A