AI dialogue interaction method and device, storage medium and electronic equipment
By recognizing multimodal trigger requests and launching adapted dialogue windows, combined with device sensors and dynamically adjusting the window display, the convenience and scenario adaptability issues of existing AI dialogue interaction technologies are solved, improving response speed and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-14
AI Technical Summary
Existing AI dialogue interaction technologies lack ease of interaction and have weak scene adaptability, failing to meet users' diverse usage scenarios and real-time interaction needs, resulting in problems such as slow response, cumbersome operation, and wasted display space.
By identifying the valid trigger type of multimodal trigger requests, the corresponding temporary or regular dialog window is launched, and the window display form is dynamically adjusted based on the characteristics of the interactive content to achieve smooth switching. Combined with device sensors, the triggering method is expanded, and multiple sensor signals are integrated for accurate identification and verification.
It has improved the scene adaptability and user experience of AI dialogue interaction, increased response speed, reduced operating costs, ensured information continuity and smooth interaction, and reduced the probability of accidental triggering.
Smart Images

Figure CN121858708A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an AI dialogue interaction method, device, storage medium and electronic device. Background Technology
[0002] With the rapid development of artificial intelligence (AI) technology, AI-powered conversational interaction has been widely applied to terminal products such as shopping, office work, social networking, and mobile operating systems. It has also been deeply integrated into core medical and financial scenarios such as healthcare applications (apps), online consultation platforms, banking apps, and insurance service platforms, becoming a core interaction method for users to obtain consultations, conduct business, manage health, and connect with financial services. With the widespread adoption of mobile internet, user scenarios are becoming increasingly diverse. This places higher demands on the convenience, immediacy, and scenario adaptability of AI conversational interaction. Especially in medical scenarios, it is necessary to balance information accuracy and response timeliness, while in financial scenarios, operational security and service professionalism must be guaranteed, further driving the industry towards more flexible, efficient, and precise interaction models.
[0003] Currently, most mainstream AI dialogue interaction technologies adopt a single entry point, fixed process, and fixed window design. In terms of interaction triggering, the entry point is mostly a separate button, requiring users to go through a fixed process of "click → load page → trigger interaction." Products supporting voice interaction also require the entry point to be activated first or voice to be activated separately, failing to achieve direct linkage between the entry point and voice, nor to expand triggering methods by incorporating device sensors. Dialogue windows are mostly of a fixed size, with only a few supporting manual adjustment, lacking a mechanism for automatic adaptation based on content length.
[0004] The core problems with existing technologies are insufficient ease of interaction and weak scene adaptability. Fixed processes result in slow response and cumbersome operation in high-frequency scenarios; single triggering methods only adapt to conventional scenarios where hands are free, failing to meet special needs such as when hands are busy; fixed windows either waste display space or require scrolling, increasing operational costs. These problems collectively prevent existing AI dialogue interaction technologies from fully matching users' diverse usage scenarios and real-time interaction needs, limiting the efficiency and user experience of AI dialogue functions. Summary of the Invention
[0005] In view of this, this application provides an AI dialogue interaction method, device, storage medium and electronic device, which can solve the technical problem that existing AI dialogue interaction technologies cannot fully match the diverse usage scenarios and real-time interaction needs of users, thus limiting the efficiency of AI dialogue functions and user experience.
[0006] According to a first aspect of this application, an AI dialogue interaction method is provided, comprising: In response to receiving a multimodal trigger request initiated by a user, the system identifies the valid trigger type corresponding to the multimodal trigger request, wherein the valid trigger type includes fast interaction triggers and deep interaction triggers; The corresponding target dialog window is launched according to the effective trigger type, wherein the quick interaction type trigger launches a temporary dialog window, and the deep interaction type trigger launches a regular dialog window; Based on the interactive content characteristics generated during the dialogue, the display form of the target dialogue window is dynamically adjusted; If the target dialog window is a temporary dialog window, then when the preset switching conditions are met, data is synchronously exchanged between the temporary dialog window and the regular dialog window to complete a smooth switch.
[0007] According to a second aspect of this application, an AI dialogue interaction device is provided, comprising: The identification module is used to identify the valid trigger type corresponding to the multimodal trigger request in response to receiving a multimodal trigger request initiated by the user. The valid trigger type includes fast interaction trigger and deep interaction trigger. The startup module is used to launch the corresponding target dialog window according to the effective trigger type, wherein the quick interaction type trigger launches a temporary dialog window, and the deep interaction type trigger launches a regular dialog window; The adjustment module is used to dynamically adjust the display form of the target dialogue window based on the interactive content features generated during the dialogue process; The switching module is used to synchronously exchange data between the temporary dialog window and the regular dialog window and complete a smooth switch when the preset switching conditions are met if the target dialog window is a temporary dialog window.
[0008] According to a third aspect of this application, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the above-described AI dialogue interaction method.
[0009] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described AI dialogue interaction method.
[0010] By employing the aforementioned technical solutions, the AI dialogue interaction method, device, storage medium, and electronic device provided in this application, firstly, by responding to multimodal trigger requests and identifying effective trigger types such as fast interaction and deep interaction, can break through the limitations of the single entry point in existing technologies, simplifying the traditional lengthy process into a one-step trigger, thereby improving the response speed and scene coverage of high-frequency scenarios; secondly, by accurately launching temporary or regular dialogue windows according to the effective trigger type, it allows both fast and simple interactions and complex multi-turn interactions to have their own suitable carriers, avoiding the problem that a single window format cannot meet different interaction needs, and further enhancing scene adaptability; thirdly, it dynamically adjusts the target based on the characteristics of the interaction content. The standard dialog window display format uses adaptive adjustment instead of fixed size or manual stretching, which solves the display pain points of wasting space for short content and requiring scrolling to view long content. It also reduces user operation costs and improves the smoothness of interaction. Finally, when the temporary dialog window meets the preset switching conditions, it achieves a smooth switch with the regular dialog window by synchronizing interaction data. This ensures the continuity of information in the interaction process, retains the convenience of fast interaction, and supports the coherence of deep interaction. Ultimately, it can comprehensively improve the efficiency of AI dialogue interaction, scene adaptability, and user experience. At the same time, by effectively identifying and verifying the trigger type, it balances convenience and accuracy and reduces the probability of false triggers.
[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an AI dialogue interaction method provided in an embodiment of this application is shown; Figure 2 A flowchart illustrating an AI dialogue interaction method according to another embodiment of this application is shown; Figure 3 A schematic diagram of the structure of an AI dialogue interaction device provided in an embodiment of this application is shown. Detailed Implementation
[0013] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0014] Currently, most mainstream AI dialogue interaction technologies adopt a single entry point, fixed process, and fixed window design. In terms of interaction triggering, the entry point is mostly a separate button, requiring users to go through a fixed process of "click → load page → trigger interaction." Products supporting voice interaction also require the entry point to be activated first or voice to be activated separately, failing to achieve direct linkage between the entry point and voice, nor to expand triggering methods by incorporating device sensors. Dialogue windows are mostly of a fixed size, with only a few supporting manual adjustment, lacking a mechanism for automatic adaptation based on content length.
[0015] The core problems with existing technologies are insufficient ease of interaction and weak scene adaptability. Fixed processes result in slow response and cumbersome operation in high-frequency scenarios; single triggering methods only adapt to conventional scenarios where hands are free, failing to meet special needs such as when hands are busy; fixed windows either waste display space or require scrolling, increasing operational costs. These problems collectively prevent existing AI dialogue interaction technologies from fully matching users' diverse usage scenarios and real-time interaction needs, limiting the efficiency and user experience of AI dialogue functions.
[0016] To address the aforementioned technical problems, embodiments of the present invention provide an AI dialogue interaction method, such as... Figure 1 As shown, the method includes: Step 110: In response to receiving a multimodal trigger request initiated by the user, identify the valid trigger type corresponding to the multimodal trigger request. Valid trigger types include fast interaction triggers and deep interaction triggers.
[0017] Among them, multimodal trigger requests refer to AI dialogue interaction requests initiated by users through various different interaction forms, covering multiple trigger forms and not a single interaction method; valid trigger types refer to trigger types that have been verified for validity, confirmed to conform to the user's true interaction intent, and exclude invalid interference. They are divided into two categories according to the interaction requirement attributes: fast interaction triggers and deep interaction triggers; fast interaction triggers are trigger forms adapted to simple, immediate, and low-complexity interaction requirements, used to quickly start lightweight interaction processes; deep interaction triggers are trigger forms adapted to complex, multi-turn, and high-complexity interaction requirements, used to start full-function interaction processes.
[0018] In the embodiments of this disclosure, after receiving a multimodal trigger request from a user that covers multiple interaction forms, the AI dialogue interaction system will verify the validity of the request, eliminate meaningless interference triggers, and then accurately identify the valid trigger type that conforms to the user's true interaction intent. Based on the complexity and immediacy of the interaction needs, the valid trigger type will be further divided into fast interaction triggers that are suitable for quick and simple interactions and deep interaction triggers that are suitable for complex multi-turn interactions, providing a basis for matching the corresponding interaction window in the future.
[0019] This technology supports multimodal triggering, breaking the limitations of a single triggering method. By recognizing and classifying effective trigger types, it can accurately adapt to interaction needs of varying complexity. This ensures immediate response in fast-paced interaction scenarios and lays the foundation for deep interaction scenarios, effectively improving the scenario adaptability and trigger accuracy of AI dialogue interaction, reducing experience loss caused by invalid triggers, making the overall interaction more in line with actual user needs, and optimizing the user experience.
[0020] Step 120: Launch the corresponding target dialog window according to the valid trigger type. Among them, quick interaction triggers launch temporary dialog windows, and deep interaction triggers launch regular dialog windows.
[0021] The temporary dialogue window is a lightweight window form corresponding to quick interaction triggers, with a simple interactive interface and core functions adapted to quick interactions, used to carry short and fast dialogue content; the regular dialogue window is a full-featured window form corresponding to deep interaction triggers, with complete interactive modules and extended functions, used to carry complex multi-turn dialogue content; the target dialogue window is the dialogue window that is finally launched after accurate matching based on the effective trigger type, and is adapted to the user's interaction needs (i.e., temporary dialogue window or regular dialogue window).
[0022] In this embodiment of the disclosure, after identifying and validating the multimodal trigger request and determining its corresponding valid trigger type, the corresponding target dialogue window can be automatically associated based on a preset "trigger type-dialogue window" mapping rule. Specifically, quick interaction triggers directly point to and launch a temporary dialogue window adapted for quick and simple interaction, while deep interaction triggers point to and launch a regular dialogue window adapted for complex multi-turn interaction. This enables precise matching of trigger requirements and window functions, laying the foundation for subsequent targeted interactive content.
[0023] This technical step establishes a precise mapping mechanism between effective trigger types and target dialogue windows, allowing interaction needs of varying complexity to correspond to suitable window forms. It enables lightweight startup and efficient response in fast interaction scenarios, while ensuring full functionality support in deep interaction scenarios. It effectively balances the convenience and functionality of interaction, significantly improving the adaptability of AI dialogue interaction to different user needs, reducing unnecessary functional redundancy or missing functions that lead to experience loss, and optimizing the overall rationality of the interaction process and user experience.
[0024] Step 130: Based on the interactive content characteristics generated during the dialogue, dynamically adjust the display form of the target dialogue window.
[0025] Interactive content characteristics refer to the attributes of various interactive information generated during the dialogue process, including but not limited to core attributes such as content length, number of characters, formatting requirements, and information type (text, multimedia, etc.); dynamic adjustment refers to continuously and automatically adjusting the window-related parameters based on the real-time changing interactive content characteristics, without requiring manual operation by the user, ensuring that the adjustment action responds synchronously with the content changes; display form refers to the external appearance of the dialogue window on the interface, mainly including visually perceptible parameters such as the window's height, width, and layout structure.
[0026] In this embodiment of the disclosure, during the AI dialogue interaction, the system can capture various interactive content generated in real time, extract its core features such as length, number of characters, layout requirements, and information type, and automatically calculate the relevant parameters required for window adaptation display based on these real-time acquired interactive content features. Then, it continuously and dynamically adjusts the display form of the target dialogue window, such as height, width, or layout structure, to ensure that the interactive content can be fully presented in an adapted form.
[0027] This technology captures interactive content features in real time and dynamically adjusts the display form of the target dialogue window, enabling precise adaptation between the window and the interactive content. This effectively avoids problems such as content obstruction and wasted space that may occur under a fixed display form. It eliminates the need for users to manually adjust the window, reducing operational costs and making the interaction process smoother and more natural, significantly improving the adaptability of AI dialogue interaction and the user experience.
[0028] Step 140: If the target dialog window is a temporary dialog window, then when the preset switching conditions are met, the data is synchronously exchanged between the temporary dialog window and the regular dialog window and a smooth switch is completed.
[0029] Among them, the preset switching conditions refer to the pre-set criteria for triggering the transition from the temporary dialog window to the regular dialog window, which is the core basis for starting the switching process; synchronous interactive data refers to the operation of completely transferring data such as dialogue records, user input information, and interaction status generated in the temporary dialog window to the regular dialog window; smooth switching refers to the switching method that maintains a natural and smooth interface transition without lag and without interrupting the user's interactive experience during the transition between the temporary dialog window and the regular dialog window.
[0030] In this embodiment of the present disclosure, when the system confirms that the target dialog window that is currently launched is a temporary dialog window, it will continuously monitor whether the preset switching conditions are met. Once the conditions are met, the data synchronization mechanism will be automatically started to completely transfer all interactive data generated in the temporary dialog window to the regular dialog window. At the same time, the switching from the temporary dialog window to the regular dialog window will be completed in a natural and smooth interface transition mode to ensure that the interactive state is not interrupted during the switching process.
[0031] This technology enables precise linkage between temporary and regular dialogue windows through preset switching conditions. Combined with the synchronous transmission of interactive data, it retains the advantage of rapid response of temporary windows while ensuring the complete functional support required for deep interaction. It avoids information loss and interaction gaps during the switching process. At the same time, the smooth switching design reduces the user's perception cost and allows for seamless connection of interaction needs of different complexities, which can significantly improve the coherence and practicality of AI dialogue interaction.
[0032] In summary, the AI dialogue interaction method provided in this application, firstly, by responding to multimodal trigger requests and identifying effective trigger types such as fast and deep interactions, breaks through the limitations of existing technologies with a single entry point, simplifying the traditional lengthy process into a one-step trigger, thus improving the response speed and scene coverage in high-frequency scenarios; secondly, by accurately launching temporary or regular dialogue windows based on the effective trigger type, it allows both fast and simple interactions and complex multi-turn interactions to have their own suitable carriers, avoiding the problem that a single window format cannot accommodate different interaction needs, and further enhancing scene adaptability; thirdly, it dynamically adjusts the display form of the target dialogue window based on the characteristics of the interaction content. By using adaptive adjustment instead of fixed size or manual stretching, the display pain points of wasting space on short content and requiring scrolling to view long content can be solved, while also reducing user operation costs and improving the smoothness of interaction. Finally, when the temporary dialogue window meets the preset switching conditions, a smooth switch with the regular dialogue window is achieved by synchronizing interaction data, which can ensure the continuity of information in the interaction process, retain the convenience of fast interaction, and support the coherence of deep interaction. Ultimately, this can comprehensively improve the efficiency of AI dialogue interaction, scene adaptability, and user experience. At the same time, by effectively identifying and verifying the trigger type, convenience and accuracy are balanced, and the probability of false triggers is reduced.
[0033] Furthermore, as a refinement and extension of the specific implementation methods of the above embodiments, and to fully illustrate the implementation methods of this embodiment, this embodiment also provides another AI dialogue interaction method, such as... Figure 2 As shown, the method includes: Step 210: In response to receiving a multimodal trigger request initiated by the user, receive in real time the multimodal interaction signals collected by the sensors configured by the user equipment, including accelerometers, vibration sensors, touch sensors, acoustic sensors and image acquisition sensors.
[0034] Among them, multimodal interaction signals refer to the physical signals collected by various sensors of the device corresponding to multimodal trigger requests. Different trigger forms correspond to different types of sensor signals, which are the core data basis for identifying the user's trigger intention. Accelerometers are sensors configured on user devices to monitor three-dimensional acceleration changes. They can capture acceleration data generated by the shaking and movement of objects, providing signal support for gesture triggers such as shaking. Vibration sensors are sensors used to monitor the vibration state of the device. They can identify characteristic vibration waveforms generated by knocking, collision, etc., and are compatible with signal acquisition for trigger methods such as double-tapping the device back panel. Touch sensors are sensors integrated on the device's touch screen. They can monitor the position, trajectory, and movement changes of multi-finger touch, providing signal acquisition support for touch triggers such as three-finger swipe down. Acoustic sensors are sensors used to collect sound signals (such as MEMS acoustic sensors). They can capture user voice information and provide sound data for voice wake-up or gaze + voice linkage wake-up. Image acquisition sensors refer to the device's front-facing camera or infrared sensor, which can detect visual information such as the user's gaze state in real time, providing image data support for trigger methods such as gaze + voice linkage.
[0035] In the embodiments of this disclosure, when the system responds to a multimodal trigger request initiated by the user that covers multiple interaction forms, it will simultaneously call the sensor devices such as the accelerometer, vibration sensor, touch sensor, acoustic sensor and image acquisition sensor pre-configured on the device to capture the physical signals corresponding to the user's trigger action in real time. These signals from different sensors together constitute multimodal interaction signals, which can provide comprehensive data support for subsequent identification of valid trigger types and judgment of user interaction intentions.
[0036] This technology integrates multiple types of sensors on the device and collects multimodal interaction signals in real time, breaking the limitations of single sensor signal acquisition. It can comprehensively capture the core data corresponding to different triggering methods, ensuring the effective implementation of the multimodal triggering system and laying a solid data foundation for subsequent accurate identification of user triggering intentions and realization of scenario-based interaction. It can significantly improve the scenario adaptability of AI dialogue interaction and the comprehensiveness and accuracy of trigger recognition.
[0037] Step 220: Extract the core feature parameters of the multimodal interaction signal. The core feature parameters include at least two of the following: action duration, signal amplitude, sliding trajectory, speech signal-to-noise ratio, and gaze duration.
[0038] Among them, the core feature parameters refer to the key quantitative indicators extracted from multimodal interaction signals that can characterize the effectiveness of the triggered action and the user's interaction intention, and are the core basis for determining whether the trigger type is effective; the action duration refers to the duration of a single triggered action or the action interval time (such as the sliding time of a three-finger swipe down, the interval between two taps of double-tapping the back panel), used to distinguish between effective triggers and unintentional operations; the signal amplitude refers to the intensity quantification value of the multimodal interaction signal (such as the shaking intensity detected by the accelerometer, the tapping vibration amplitude captured by the vibration sensor), and is a key indicator for determining whether the triggered action meets the effective standard; the sliding trajectory refers to the path of the finger's position change on the screen when the user triggers it through touch (such as a three-finger swipe down), used to accurately identify the effectiveness of touch-based triggered actions; the speech signal-to-noise ratio refers to the ratio of effective speech to environmental noise in the wake-up speech signal collected by the acoustic sensor, reflecting the clarity of the speech signal, and is the core parameter for determining the effectiveness of voice-based triggers; the gaze duration refers to the time the user continuously gazes at the device screen as detected by the image acquisition sensor (front camera / infrared sensor), used to verify the user's interaction intention in gaze + voice linkage triggering.
[0039] In the embodiments of this disclosure, after the system acquires multimodal interaction signals collected by various sensors of the device corresponding to different triggering methods, it can use a specific algorithm to filter and extract key information that can accurately reflect the attributes of the triggering action and the user's intention from these signals. That is, it includes at least two core feature parameters among the action duration, signal amplitude, sliding trajectory, speech signal-to-noise ratio, and gaze duration. This can provide quantifiable core data support for subsequent judgment of the validity of the triggering type and identification of user interaction needs.
[0040] By extracting the core feature parameters of multimodal interaction signals, objective and accurate quantitative evidence can be provided for determining the validity of multimodal triggers. This effectively distinguishes between valid triggers and unintentional operations or environmental interference, significantly improving the accuracy and reliability of multimodal trigger recognition. Furthermore, by adapting to the differentiated judgment requirements of different triggering methods, a data foundation can be laid for the implementation of anti-accidental touch mechanisms, avoiding interference from invalid triggers on the user experience and ensuring the stable operation and efficient response of the multimodal interaction system.
[0041] Step 230: Match the core feature parameters with the preset trigger type feature library to determine the trigger type corresponding to the multimodal interaction signal.
[0042] The preset trigger type feature library pre-stores the feature parameter ranges for fast interactive triggers and deep interactive triggers.
[0043] In this embodiment of the present disclosure, after extracting core feature parameters such as action duration, signal amplitude, and sliding trajectory from the multimodal interaction signal, the system will compare these real-time extracted parameters with the feature parameter ranges of fast interaction type triggers and deep interaction type triggers pre-stored in the preset trigger type feature library. By judging the parameter range to which the core feature parameter belongs, the trigger type corresponding to the current multimodal interaction signal is accurately determined, thus completing the coherent process from data extraction to type determination.
[0044] By matching core feature parameters with a preset trigger type feature library, the trigger type corresponding to multimodal interaction signals can be accurately determined, clearly distinguishing the different needs of rapid interaction and deep interaction. This provides a reliable basis for subsequently launching adapted dialogue windows, effectively avoiding confusion and misjudgment of different trigger types, improving the stability and accuracy of the multimodal trigger system, and making the determination of trigger types systematic, ensuring the orderly connection of subsequent interaction processes, and optimizing the accuracy and scene adaptation efficiency of AI dialogue interaction as a whole.
[0045] Step 240: For the identified trigger types, call the corresponding validity verification rules to verify the multimodal interaction signals and filter out the valid trigger types.
[0046] For embodiments of this disclosure, step 240 may include the following steps: Step 240-1: Retrieve the verification threshold corresponding to the trigger type. The verification threshold for quick interaction triggers includes the action parameter threshold and the scene adaptation threshold. The verification threshold for deep interaction triggers includes the operation confirmation threshold.
[0047] Among them, the verification threshold is a pre-set set of standard parameters used to determine whether the trigger type matches the user's true intention, and it is the core basis for screening valid triggers and filtering invalid interference; the action parameter threshold is a quantitative standard set for the action attributes of quick interaction triggers (such as the peak acceleration of shaking, the tapping interval and vibration amplitude of double-tapping the back panel, the sliding distance and time of three-finger swipe, etc.), used to distinguish between valid trigger actions and unintentional operations and environmental interference; the scene adaptation threshold is an auxiliary judgment standard for quick interaction triggers set in combination with the user's usage scenario (such as the voice signal-to-noise ratio and gaze duration requirements in different scenarios, etc.), which further improves the accuracy of triggers through scenario-based intelligent judgment; the operation confirmation threshold is a judgment standard set for deep interaction triggers, used to verify the user's clear operation intention (such as the pressing duration of the click action, the accuracy of the click position, etc.), to avoid accidental operation to start complex interaction processes.
[0048] In this embodiment of the disclosure, after determining the trigger type (fast interaction trigger or deep interaction trigger) corresponding to the multimodal interaction signal by matching the core feature parameters with the preset trigger type feature library, the system will automatically retrieve the verification threshold pre-bound to the trigger type. For fast interaction triggers, a combination judgment standard including action parameter threshold and scene adaptation threshold is retrieved, and for deep interaction triggers, a dedicated operation confirmation threshold is retrieved. This can provide clear and targeted judgment basis for subsequent accurate verification of trigger validity and confirmation of user interaction intent.
[0049] By configuring dedicated verification thresholds for fast-interaction triggers and deep-interaction triggers respectively, fast-interaction triggers achieve precise prevention of accidental touches through the synergy of action parameter thresholds and scene adaptation thresholds, while deep-interaction triggers ensure the clarity of the launch intent through operation confirmation thresholds. This approach can meet the convenient response needs in fast-interaction scenarios while avoiding the risk of misoperation in deep-interaction scenarios, effectively improving the accuracy and stability of multimodal AI dialogue interaction, achieving a balance between convenience and reliability, and optimizing the user experience in different interaction scenarios.
[0050] Step 240-2: Compare the feature parameters of the multimodal interaction signal with the verification threshold to initially screen out candidate trigger signals that meet the parameter standards.
[0051] Among them, feature parameters refer to key indicators extracted from multimodal interaction signals that can quantify the core attributes of the triggering action, including action duration, signal amplitude, sliding trajectory, speech signal-to-noise ratio, gaze duration, etc., which are the core data basis for determining the validity of the triggering action; candidate triggering signals refer to triggering signals that meet the preset judgment criteria and are initially identified as having effective interaction intentions after being compared with feature parameters and verification thresholds, and are the objects of further scenario verification.
[0052] In this embodiment of the present disclosure, after extracting the feature parameters such as the action duration, signal amplitude, and sliding trajectory corresponding to the multimodal interaction signal, the system will retrieve the verification threshold bound to the determined trigger type (fast interaction trigger or deep interaction trigger), and then compare each feature parameter with the corresponding verification threshold one by one to determine whether the feature parameter is within the preset valid value range. Finally, the system will select the trigger signal that meets the verification threshold requirements and initially has a valid interaction intention as a candidate trigger signal, providing a basis for subsequent secondary scene verification.
[0053] By comparing and filtering the characteristic parameters of multimodal interaction signals with corresponding verification thresholds, invalid trigger signals can be initially filtered out, effectively eliminating unintentional triggers caused by non-standard actions or insufficient signal strength. This lays a solid foundation for accurate identification of valid trigger types, significantly improves the initial judgment accuracy of the multimodal triggering system, reduces the processing pressure of subsequent verification, and ensures the efficiency of interactive response. This allows trigger judgment to be supported by quantitative standards and to avoid some invalid interference in advance, thus optimizing the overall reliability and smoothness of the interaction.
[0054] Step 240-3: Perform scenario adaptation verification on the candidate trigger signals based on the current usage scenario of the user equipment.
[0055] For embodiments of this disclosure, the steps may include: real-time acquisition of multi-dimensional scene data of user devices, the multi-dimensional scene data including at least environmental feature data, device status feature data and user behavior feature data; determining the current usage scenario corresponding to the multi-dimensional scene data based on preset scene classification rules, the current usage scenario including any one of low threshold triggering scenario, high threshold triggering scenario and prohibited triggering scenario; and performing scene adaptation verification on candidate triggering signals based on preset verification rules of the current usage scenario.
[0056] Among them, scenario adaptation verification combines the attributes of the user device's current usage scenario to perform targeted validity verification on candidate trigger signals, ensuring that the trigger action matches the scenario requirements and filtering out false trigger signals that do not match the scenario; multi-dimensional scenario data is a comprehensive dataset used to support scenario determination, which at least includes environmental feature data reflecting the surrounding environment, device status feature data representing the device's own operating status, and user behavior feature data reflecting the user's operational behavior patterns; preset scenario classification rules are pre-defined scenario division standards, a rule system that clearly classifies usage scenarios into different trigger control types based on the analysis results of multi-dimensional scenario data; low Threshold triggering scenarios are designed for situations where users are busy and require immediate interaction. In these scenarios, the trigger verification standards are relatively lenient, prioritizing the convenience and immediacy of the interaction. High threshold triggering scenarios are sensitive to environmental interference and require strict control over accidental triggering. In these scenarios, the trigger verification standards are more stringent, with the core focus on ensuring the accuracy of the triggering action. Prohibited triggering scenarios refer to scenarios where AI dialogue interaction is not required. In these scenarios, the trigger response will be directly blocked to avoid interference from invalid interactions. Preset verification rules are dedicated trigger verification logics pre-defined for different categories of scenarios, clarifying the verification standards and judgment process for candidate trigger signals in each scenario, and achieving scenario-based precise verification.
[0057] By collecting multi-dimensional scene data and classifying scenes based on preset rules, and then executing exclusive preset verification rules for different scenes, scenario-based accurate verification of candidate trigger signals can be achieved. This not only ensures the convenience and immediacy of triggering interactions in the adapted scenarios, but also effectively filters out false trigger signals that do not match the scenarios through high threshold verification mechanisms and trigger prohibition mechanisms. This significantly improves the triggering accuracy and scene adaptability of the multimodal triggering system, balances the convenience of interaction and the reliability of use, and further optimizes the overall user experience of AI dialogue interaction from the perspective of scene adaptation.
[0058] Step 240-4: Determine the trigger types that pass the parameter comparison and scene adaptation verification as valid trigger types.
[0059] Step 250: Launch the corresponding target dialog window according to the valid trigger type. Among them, quick interaction triggers launch temporary dialog windows, and deep interaction triggers launch regular dialog windows.
[0060] For embodiments of this disclosure, step 250 may include the following steps: Step 250-1: Based on the preset scene mapping relationship between trigger types and dialog windows, determine the target dialog window corresponding to the valid trigger type.
[0061] Among them, the preset scene mapping relationship between trigger types and dialogue windows refers to the pre-set and stored rule system that binds different trigger types with corresponding adapted dialogue windows, clearly defining the target dialogue windows corresponding to fast interaction triggers and deep interaction triggers, which can provide a clear basis for the accurate matching of trigger types and windows.
[0062] In specific application scenarios, the system will pre-establish and store scene mapping rules between trigger types and dialog windows, clarifying the binding relationship between quick interaction triggers and corresponding temporary dialog windows, and deep interaction triggers and corresponding regular dialog windows. After completing the validity verification of multimodal trigger requests and determining the valid trigger type, the system can automatically match and lock the dialog window that is compatible with the current valid trigger type based on the preset mapping relationship, and use it as the target dialog window to carry subsequent interactions.
[0063] By pre-setting the scene mapping relationship between trigger types and dialogue windows, it is possible to achieve accurate and rapid matching between effective trigger types and target dialogue windows. This allows different interaction needs (quick and simple or complex and multi-turn) to be directly matched with the appropriate window form, avoiding the problem that a single window cannot take into account both convenience and functionality. It also saves users from manually selecting window types, shortens the interaction startup path, improves the scene adaptation accuracy and startup efficiency of AI dialogue interaction, and makes the overall interaction process more in line with user needs, thus optimizing the user experience.
[0064] Step 250-2: Retrieve the initial configuration parameters of the target dialog window.
[0065] Among them, the initial configuration parameters of the window refer to the basic running parameters and interface configuration information that are pre-set for different types of target dialog windows and loaded by default when the window starts. They are the basis for the initial form and core functions of the window after it starts, including but not limited to key configuration items such as interface layout, default function modules, display style, and interaction permissions.
[0066] In this embodiment of the disclosure, after determining the target dialogue window (temporary dialogue window or regular dialogue window) corresponding to the current effective trigger type through scene mapping relationship, the system will automatically retrieve and call the pre-stored initial configuration parameters specific to this type of window. The initial configuration parameters of the temporary dialogue window focus on lightweight and immediacy (such as default activation of voice interaction, simplified interface layout, setting the initial pop-up position, etc.), while the initial configuration parameters of the regular dialogue window focus on full functionality and scalability (such as loading the complete interaction module, reserving a history area, configuring multiple forms of input permissions, etc.), laying the foundation for the window to start quickly and adapt to the corresponding interaction scenario.
[0067] By retrieving the specific initial configuration parameters corresponding to the target dialogue window, the window starts with the basic form and core functions adapted to the corresponding interaction scenario, without the need for subsequent additional configuration adjustments. This ensures both the efficiency of window startup and the accurate matching of the initial state with the interaction requirements, avoiding the impact of functional redundancy or missing functions on the interactive experience. At the same time, the standardized initial configuration parameters make the interaction startup process more standardized in different scenarios, which can improve the stability and scene adaptation accuracy of AI dialogue interaction and optimize the overall smoothness of interaction.
[0068] Step 250-3: Launch the target dialog window at the preset page position based on the initial window configuration parameters.
[0069] The preset page position refers to a fixed launch display area pre-defined according to the window type, ensuring that the window adapts to the scenario requirements after launch and does not affect user operation. It is divided into two categories: a dedicated area for temporary dialog windows and a dedicated area for regular dialog windows. When the target dialog window is a temporary dialog window, the preset page position is the top center or bottom pop-up area of the current device's primary page, with a preset gap between the pop-up edge and the screen edge to avoid obscuring the core function entry point. When the target dialog window is a regular dialog window, the preset page position is displayed as an independent new page within the current application hierarchy, retaining the original page's return entry at the top. The window page must at least include a history area, an input area, and a function area. The primary page refers to the core basic page (such as the homepage, function category homepage, etc.) that can be directly accessed without navigating, serving as the upper-level launch platform for the temporary dialog window. The preset gap refers to a fixed distance pre-defined between the edge of the temporary dialog window and the edge of the device screen, used to ensure interface aesthetics and operational safety. The core function entry point refers to the key operation entry point frequently used by users on the first-level page (such as core function buttons, frequently used service entry points, etc.), which should be avoided when launching a temporary dialog window; the independent new page form refers to a completely new page form generated independently in the application layer, detached from the current first-level page, with independent interface layout and functional space; the history area is the functional area in the regular dialog window used to store and display past interaction content (user input, AI replies, etc.); the input area is the functional area used to carry user input operations (text, voice, etc.), which is the core operation entry point for user interaction with AI; the function area is the area of auxiliary interactive function modules integrated in the regular dialog window (such as function selection, settings entry points, etc.), which can provide extended support for complex interactions.
[0070] In this embodiment of the disclosure, after retrieving the initial configuration parameters of the target dialog window, the system can launch the window at a preset page position according to the type of the target dialog window (temporary dialog window or regular dialog window): if it is a temporary dialog window, it is displayed in the upper center or bottom pop-up area of the current device's first-level page, while ensuring that the edge of the pop-up window and the edge of the screen are reserved with a preset distance to avoid obscuring the core function entrance of the page; if it is a regular dialog window, it is displayed in the current application level as an independent new page, the original page return entrance is retained at the top of the page, and the window page loads the history area, input area and function area and other core modules according to the initial configuration parameters to complete the window launch and adaptation.
[0071] By launching the target dialogue window at a dedicated preset page position based on the initial window configuration parameters, the temporary dialogue window can quickly carry out simple interactions without obscuring the core functions, while the regular dialogue window provides complete interaction support as an independent page. This ensures both the immediacy and ease of operation of rapid interactions, as well as the functional integrity and process continuity of complex interactions. At the same time, the fixed preset position and unified initial configuration make window launch more standardized and user operation expectations clearer, which can effectively reduce the interaction learning cost and improve the scenario adaptability and overall user experience of AI dialogue interaction.
[0072] Step 260: Based on the interactive content characteristics generated during the dialogue, dynamically adjust the display form of the target dialogue window.
[0073] For embodiments of this disclosure, step 260 may include the following steps: Step 260-1: Collect interactive content during the dialogue in real time. The interactive content includes user speech-to-text content, AI reply text content, and multimedia interactive elements.
[0074] Interactive content refers to the collection of all information related to user interaction generated during AI dialogue; user speech-to-text content refers to the recognizable and displayable text data generated by the system after the user triggers interaction through voice, which is one of the core forms of carrying user interaction intentions; AI response text content refers to the text-based response information generated by the system in response to the user's requests and instructions, which is the core carrier for conveying interaction results and realizing information exchange; multimedia interactive elements refer to various elements other than text used to assist interaction and enrich the form of information transmission during the dialogue, including but not limited to images, voice clips, file attachments, emoticons, etc.
[0075] In this embodiment of the disclosure, after the AI dialogue interaction is initiated, the system will activate a real-time information capture mechanism to continuously and synchronously collect various key information generated during the interaction process. This includes text content generated by speech-to-text processing after the user triggers the interaction through voice, as well as AI response text content from the system in response to user needs. It will also capture multimedia interactive elements such as images and voice clips involved in the dialogue, and comprehensively collect various interactive data related to the dialogue.
[0076] By collecting user speech-to-text content, AI response text content, and multimedia interactive elements during the dialogue process in real time, it can provide complete and accurate data support for functions such as adaptive adjustment of the dialogue window and data synchronization between different windows. This ensures that the adaptive window can dynamically match the display requirements of the interactive content, while guaranteeing the integrity and continuity of interactive information when switching windows, enriching the forms of interactive information, making AI dialogue interaction more accurate, comprehensive, and smooth, and significantly improving information transmission efficiency and user experience.
[0077] Step 260-2: Analyze the length, number of characters, layout requirements, and element types of the interactive content, and calculate the corresponding minimum display space.
[0078] Among them, the length of interactive content refers to the extension dimension of interactive content when presented on the interface, including quantifiable spatial attributes such as the extension of text lines and the size extension of multimedia elements; the number of characters refers to the total number of characters in text-based information (user speech-to-text, AI reply text) in interactive content, which is the core quantitative indicator for calculating text display space; layout requirements refer to the display format requirements of interactive content to ensure readability, including text wrapping rules, paragraph spacing, and the layout relationship between multimedia and text; element type refers to the classification of the carrier form of interactive content, mainly divided into text elements and multimedia elements such as images and audio clips, and the display space requirements of different types of elements are different; minimum display space refers to the minimum interface area that the window needs to occupy under the premise of meeting layout requirements and ensuring the complete presentation of interactive content, which is the core reference standard for dynamic window adjustment.
[0079] In this embodiment of the disclosure, after the system collects the interactive content (including text, multimedia and other elements) during the dialogue process in real time, it can comprehensively analyze the key attributes of the content, clarify the length and number of characters of the text content, the format requirements required for the overall layout, and the element type corresponding to each content. Then, combined with the display rules and layout specifications of different element types, it can comprehensively calculate the minimum interface area that can fully carry the current interactive content without generating redundant space, that is, the minimum display space, so as to provide a precise basis for the dynamic adaptation and adjustment of the target dialogue window.
[0080] By analyzing the core attributes of interactive content and calculating the minimum display space, the dynamic adjustment of the target dialogue window has a precise quantitative reference, ensuring that the window can present the interactive content completely in the most compact and reasonable form. This avoids the space waste or content obstruction problems caused by fixed windows, and eliminates the need for manual user intervention. It can significantly improve the adaptability and rationality of window display, making the interactive interface more concise and intuitive, while ensuring the readability of different types of interactive content, and optimizing the overall smoothness of AI dialogue interaction and user experience.
[0081] Step 260-3: If the target dialog window is a temporary dialog window, the window height is dynamically adjusted based on the minimum display space so that the adjusted window height does not exceed the preset height percentage of the device screen height, and the interactive content is scrolled and displayed within the temporary dialog window after the window height is adjusted.
[0082] Among them, the preset height ratio refers to the pre-set maximum height of the temporary dialog window and the ratio threshold of the device screen height, which is used to limit the maximum height of the window and prevent the window from occupying too much screen space; scrolling display refers to the window enabling a scrolling mechanism when the total amount of interactive content exceeds the display range of the adjusted temporary dialog window, allowing users to view the display of the complete interactive content by swiping.
[0083] In this embodiment of the disclosure, when the target dialog window is determined to be a temporary dialog window, the system can automatically adjust the window height in real time based on the previously calculated minimum display space of the interactive content to ensure that the window can fully accommodate the current interactive content. At the same time, the adjusted window height is strictly controlled to not exceed the preset percentage of the device screen height. If the total amount of interactive content exceeds the display range of the adjusted window, the scrolling display function is enabled within the window to ensure that the user can view all interactive content.
[0084] By dynamically adjusting the height of temporary dialog windows based on the minimum display space and limiting the maximum size of the window with a preset height percentage, it can achieve precise adaptation between the window and the interactive content, avoiding space waste or content obstruction caused by fixed sizes. It can also prevent the window from occupying too much screen space and affecting other operations. Combined with the scrolling display function, it ensures the complete presentation of long content. No manual adjustment by the user is required, which can significantly reduce the operating cost, make the interface display more reasonable and the interaction smoother in fast interaction scenarios, and comprehensively optimize the user experience.
[0085] Step 260-4: If the target dialog window is a regular dialog window, then adaptively adjust the window height or width based on the minimum display space so that the adjusted window height supports the complete display of interactive content.
[0086] In this embodiment of the present disclosure, once the target dialog window is determined to be a regular dialog window, the system can use the minimum display space calculated after parsing the interactive content (including text, multimedia, etc.) as a benchmark to automatically start the size adaptive adjustment mechanism. Based on the length, number of characters, element type and layout requirements of the interactive content, the height or width of the window can be flexibly adjusted to ensure that the adjusted window can completely and clearly carry all interactive content. The precise adaptation of content and window form can be achieved without manual operation by the user.
[0087] By adaptively adjusting the height or width of a regular dialogue window based on the minimum display space, problems such as content occlusion and space waste that easily occur in complex multi-turn interactions with fixed-size windows can be effectively solved. This ensures the complete presentation of various interactive contents without requiring users to manually adjust the window size, reducing operational costs. At the same time, it makes the interface display in complex interaction processes more in line with content needs, improves the smoothness and usability of interaction, and fully guarantees a good user experience in deep dialogue scenarios.
[0088] Step 270: If the target dialog window is a temporary dialog window, then when the preset switching conditions are met, the data is synchronously exchanged between the temporary dialog window and the regular dialog window and a smooth switch is completed.
[0089] The preset switching conditions include at least one of the following: the adjusted window height reaches a preset height percentage of the device screen height; the length of the interactive content exceeds a preset number of characters; the user actively initiates a switching command; or multiple rounds of interaction triggering.
[0090] For embodiments of this disclosure, step 270 may include the following steps: Step 270-1: Monitor the height changes, content length, and user operation commands of the temporary dialog window in real time to determine whether the preset switching conditions are met.
[0091] In this embodiment of the present disclosure, during the operation of the temporary dialog window, the system can continuously and synchronously monitor the dynamic changes in the vertical height of the window and the total length of the internal interactive content. At the same time, it can capture the relevant operation commands issued by the user during the interaction in real time. The system can comprehensively compare and analyze these three types of real-time collected information with the pre-set switching judgment criteria to determine whether the current conditions for the conversion from the temporary dialog window to the regular dialog window are met, thus providing an accurate basis for whether to start the window switching and data synchronization process in the future.
[0092] By monitoring the height changes, content length, and user operation commands of temporary dialogue windows in real time and determining whether preset switching conditions are met, precise and timely linkage between temporary and regular dialogue windows can be achieved. This ensures a lightweight and efficient experience in fast-paced interaction scenarios, while also allowing for timely switching to a full-featured window when content complexity increases or users have deeper interaction needs. This avoids issues such as content overflow and insufficient functionality in temporary windows, ensuring the continuity and integrity of the interaction process. It can significantly improve the adaptability of AI dialogue interaction to different complexity requirements, making the overall interaction process more in line with actual user scenarios and optimizing the user experience.
[0093] Step 270-2: When the preset switching conditions are met, generate a window switching command and freeze the current interactive state of the temporary dialog window.
[0094] Among them, the window switching instruction refers to the control signal generated by the system after the preset switching conditions are met, which is used to trigger the conversion of the temporary dialog window to the regular dialog window. It is the core instruction to start the window switching process. The current interaction state refers to the real-time running state of the temporary dialog window when the switching conditions are met, including key information such as the generated dialog records, the user's current input progress, window display parameters, and unfinished interaction instructions.
[0095] In this embodiment of the present disclosure, when the system determines that the preset switching conditions have been met by monitoring the height change, content length and user operation instructions of the temporary dialog window, it will immediately and automatically generate a window switching instruction to trigger the window switching. At the same time, it will lock and freeze the current dialogue record, input progress, display parameters and other interactive states of the temporary dialog window to prevent information loss or state disorder during the switching process. This can provide a basis for the subsequent synchronization of interactive data and smooth switching between the temporary dialog window and the regular dialog window.
[0096] By generating a window switching command and freezing the current interactive state of the temporary dialogue window when the preset switching conditions are met, the integrity and stability of the interactive information during the window switching process can be ensured, avoiding interaction gaps caused by data loss or state disorder. This provides a reliable guarantee for subsequent smooth switching and data synchronization, allowing users to continue the previous dialogue logic and operation progress when transitioning from fast interaction scenarios to deep interaction scenarios, significantly improving the continuity of AI dialogue interaction and user experience.
[0097] Step 270-3: Establish a data synchronization channel between the temporary dialog window and the regular dialog window.
[0098] In this embodiment of the present disclosure, after freezing the current interactive state of the temporary dialog window, the system will immediately start the dedicated communication link construction process to establish a data synchronization channel between the temporary dialog window and the regular dialog window. The transmission protocol and data format of the channel are specified to ensure that key data such as the dialog records, user input progress, and interactive state parameters generated in the temporary dialog window can be transmitted to the regular dialog window in real time and without omission through the channel, providing data support for the continued interaction after subsequent window switching.
[0099] By establishing a data synchronization channel between temporary and regular dialogue windows, seamless transfer of interactive data between the two types of windows can be achieved. This ensures that core information such as historical dialogue records and interaction status are not lost or deviated when switching windows, allowing users to continue the previous dialogue logic when transitioning from quick interaction to deep interaction. This avoids repetitive operations or re-entry of information, significantly improving the coherence and integrity of AI dialogue interaction and optimizing the user experience in different scenarios.
[0100] Step 270-4: Based on the data synchronization channel, smoothly switch the temporary dialog window to the regular dialog window according to the preset transition animation, and restore the interactive state after synchronously loading the interactive data in the regular dialog window.
[0101] Among them, preset transition animations refer to the interface transition effects (such as gradual enlargement, smooth stretching, seamless connection, etc.) set in advance when switching windows, which are used to optimize the visual experience of switching and reduce the perceived gaps in user operation.
[0102] In this embodiment of the present disclosure, based on the establishment of a data synchronization channel between the temporary dialog window and the regular dialog window, the system will drive the temporary dialog window to smoothly transition to the regular dialog window according to the preset transition animation effect. At the same time, through the data synchronization channel, the frozen interactive data in the temporary dialog window will be transmitted to the regular dialog window in real time and completely and loaded. After the data is loaded, the regular dialog window will automatically restore the interactive state, ensuring that the user can seamlessly connect to the previous dialog process.
[0103] By leveraging a data synchronization channel to ensure the seamless transmission of interactive data, coupled with preset transition animations to guarantee the visual smoothness of window switching, and restoring the interactive state to allow the dialogue flow to continue seamlessly, it can completely avoid information loss and interaction gaps during the switching process, and reduce the user's perception cost of window switching. This allows the transition from fast interaction to deep interaction to be natural and smooth, without requiring the user to repeat operations or re-enter, which can significantly improve the coherence, integrity and overall user experience of AI dialogue interaction.
[0104] In summary, the technical solution in this application, by responding to multimodal trigger requests and identifying effective trigger types such as fast-interaction and deep-interaction, can break the limitations of the single entry point in existing technologies, simplifying the traditional lengthy process into a one-step trigger, thereby improving the response speed and scene coverage of high-frequency scenarios. Secondly, by accurately launching temporary or regular dialogue windows based on the effective trigger type, it allows both fast and simple interactions and complex multi-turn interactions to have their own suitable carriers, avoiding the problem that a single window form cannot meet different interaction needs, and further enhancing scene adaptability. Furthermore, by dynamically adjusting the display form of the target dialogue window based on the characteristics of the interactive content, it can adapt to different scenarios. Adjusting the size instead of a fixed size or manual stretching solves the display pain points of wasting space on short content and requiring scrolling to view long content, while also reducing user operation costs and improving the smoothness of interaction. Finally, when the temporary dialogue window meets the preset switching conditions, a smooth switch with the regular dialogue window is achieved by synchronizing interaction data, which ensures the continuity of information in the interaction process, retains the convenience of fast interaction, and supports the coherence of deep interaction. Ultimately, this can comprehensively improve the efficiency of AI dialogue interaction, its scene adaptability, and user experience. At the same time, by effectively identifying and verifying the trigger type, it balances convenience and accuracy and reduces the probability of false triggers.
[0105] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides an AI dialogue interaction device, such as... Figure 3 As shown, the device includes: an identification module 31, a startup module 32, an adjustment module 33, and a switching module 34.
[0106] The identification module 31 can be used to respond to a multimodal trigger request initiated by a user and identify the valid trigger type corresponding to the multimodal trigger request. The valid trigger types include fast interaction triggers and deep interaction triggers. The startup module 32 can be used to launch the corresponding target dialog window according to the valid trigger type, wherein a quick interaction trigger launches a temporary dialog window, and a deep interaction trigger launches a regular dialog window. The adjustment module 33 can be used to dynamically adjust the display form of the target dialogue window based on the interactive content features generated during the dialogue process; The switching module 34 can be used to synchronously exchange data between the temporary dialog window and the regular dialog window and complete a smooth switch when the preset switching conditions are met if the target dialog window is a temporary dialog window.
[0107] In some embodiments of this application, the identification module 31 is specifically used to receive multimodal interaction signals collected by sensors configured on the user equipment in real time. The sensors include an accelerometer, a vibration sensor, a touch sensor, an acoustic sensor, and an image acquisition sensor. It extracts core feature parameters of the multimodal interaction signals, including at least two of the following: action duration, signal amplitude, sliding trajectory, speech signal-to-noise ratio, and gaze duration. It matches the core feature parameters with a preset trigger type feature library to determine the trigger type corresponding to the multimodal interaction signal. The preset trigger type feature library pre-stores feature parameter ranges for fast interaction triggers and deep interaction triggers. For the identified trigger type, it calls the corresponding validity verification rules to verify the multimodal interaction signal and filter out valid trigger types.
[0108] In some embodiments of this application, when verifying the multimodal interaction signal by calling the corresponding validity verification rules for the identified trigger type and filtering out valid trigger types, the identification module 31 can specifically be used to retrieve the verification threshold corresponding to the trigger type. The verification threshold for fast interaction triggers includes an action parameter threshold and a scene adaptation threshold, while the verification threshold for deep interaction triggers includes an operation confirmation threshold. The feature parameters of the multimodal interaction signal are compared with the verification threshold to initially filter out candidate trigger signals whose parameters meet the standards. Scene adaptation verification is performed on the candidate trigger signals in conjunction with the current usage scenario of the user device. Trigger types that pass both parameter comparison and scene adaptation verification are determined to be valid trigger types.
[0109] In some embodiments of this application, when performing scene adaptation verification on candidate trigger signals in conjunction with the current usage scenario of the user equipment, the identification module 31 can be specifically used to collect multi-dimensional scene data of the user equipment in real time. The multi-dimensional scene data includes at least environmental feature data, device status feature data, and user behavior feature data. Based on preset scene classification rules, the current usage scenario corresponding to the multi-dimensional scene data is determined. The current usage scenario includes any one of low threshold trigger scenario, high threshold trigger scenario, and prohibited trigger scenario. Based on preset verification rules of the current usage scenario, the candidate trigger signals are performed for scene adaptation verification.
[0110] In some embodiments of this application, the startup module 32 can be specifically used to determine the target dialog window corresponding to the effective trigger type based on the scene mapping relationship between the preset trigger type and the dialog window; retrieve the initial configuration parameters of the window corresponding to the target dialog window; and start the target dialog window at a preset page position based on the initial configuration parameters of the window. When the target dialog window is a temporary dialog window, the preset page position is the top center or bottom pop-up area of the first-level page of the current device, and the edge of the pop-up window is reserved with a preset distance from the edge of the screen so as not to obscure the core function entrance of the page. When the target dialog window is a regular dialog window, the preset page position is to jump to and display it in the current application level as an independent new page, and the original page return entrance is retained at the top of the page. The window page includes at least a history area, an input area, and a function area.
[0111] In some embodiments of this application, the adjustment module 33 can be used to collect interactive content during the dialogue process in real time. The interactive content includes user speech-to-text content, AI reply text content, and multimedia interactive elements; analyze the length, number of characters, layout requirements, and element types of the interactive content, and calculate the corresponding minimum display space; if the target dialogue window is a temporary dialogue window, the window height is dynamically adjusted based on the minimum display space so that the adjusted window height does not exceed a preset height percentage of the device screen height, and the interactive content is scrolled within the temporary dialogue window after the window height is adjusted; if the target dialogue window is a regular dialogue window, the window height or width is adaptively adjusted based on the minimum display space so that the adjusted window height supports the complete display of the interactive content.
[0112] In some embodiments of this application, the preset switching conditions include at least one of the following: the adjusted window height reaches a preset height percentage of the device screen height, the length of the interactive content exceeds a preset number of characters, the user actively initiates a switching command, and multiple rounds of interaction are triggered; the switching module 34 can be used to monitor the height change, content length, and user operation commands of the temporary dialog window in real time to determine whether the preset switching conditions are met; when the preset switching conditions are met, a window switching command is generated and the current interactive state of the temporary dialog window is frozen; a data synchronization channel is established between the temporary dialog window and the regular dialog window; based on the data synchronization channel, the temporary dialog window is smoothly switched to the regular dialog window according to a preset transition animation, and the interactive state is restored after the interactive data is synchronously loaded in the regular dialog window.
[0113] It should be noted that other corresponding descriptions of the functional units involved in the AI dialogue interaction device provided in this embodiment can be found in [reference needed]. Figure 1 and Figure 2 The corresponding descriptions in [the document] will not be repeated here.
[0114] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The AI dialogue interaction method shown.
[0115] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause an electronic device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0116] Based on the above, Figure 1 and Figure 2 The method shown, and Figure 3 To achieve the above objectives, the present application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, as shown in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 2 The AI dialogue interaction method shown.
[0117] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0118] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0119] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0120] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.
[0121] This invention, by responding to multimodal trigger requests and identifying effective trigger types such as fast and deep interactions, breaks through the limitations of existing technologies with a single entry point, simplifying the traditional lengthy process into a one-step trigger, thus improving the response speed and scene coverage in high-frequency scenarios. Secondly, by accurately launching temporary or regular dialogue windows based on the effective trigger type, it allows for both fast, simple interactions and complex, multi-turn interactions to have their own suitable platforms, avoiding the problem that a single window format cannot accommodate different interaction needs, further enhancing scene adaptability. Thirdly, by dynamically adjusting the display form of the target dialogue window based on the characteristics of the interactive content, it adaptively adjusts the display. Instead of fixed sizes or manual stretching, this approach addresses the pain points of wasted space for short content and the need for scrolling to view long content, while also reducing user operation costs and improving interaction smoothness. Finally, when a temporary dialogue window meets preset switching conditions, it achieves a smooth transition with the regular dialogue window through synchronized interaction data, ensuring the continuity of information during the interaction process. This preserves the convenience of fast interaction while supporting the coherence of deep interaction, ultimately comprehensively improving the efficiency, scene adaptability, and user experience of AI dialogue interaction. At the same time, by effectively identifying and verifying trigger types, it balances convenience and accuracy, reducing the probability of false triggers.
[0122] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0123] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. An AI dialogue interaction method, characterized in that, include: In response to receiving a multimodal trigger request initiated by a user, the system identifies the valid trigger type corresponding to the multimodal trigger request, wherein the valid trigger type includes fast interaction triggers and deep interaction triggers; The corresponding target dialog window is launched according to the effective trigger type, wherein the quick interaction type trigger launches a temporary dialog window, and the deep interaction type trigger launches a regular dialog window; Based on the interactive content characteristics generated during the dialogue, the display form of the target dialogue window is dynamically adjusted; If the target dialog window is a temporary dialog window, then when the preset switching conditions are met, data is synchronously exchanged between the temporary dialog window and the regular dialog window to complete a smooth switch.
2. The method according to claim 1, characterized in that, The identification of the valid trigger type corresponding to the multimodal trigger request includes: The system receives multimodal interaction signals in real time from sensors configured on the user equipment, including accelerometers, vibration sensors, touch sensors, acoustic sensors, and image acquisition sensors. Extract the core feature parameters of the multimodal interaction signal, which include at least two of the following: action duration, signal amplitude, sliding trajectory, speech signal-to-noise ratio, and gaze duration; The core feature parameters are matched with a preset trigger type feature library to determine the trigger type corresponding to the multimodal interaction signal. The preset trigger type feature library stores the feature parameter ranges of fast interaction type triggers and deep interaction type triggers in advance. For the identified trigger types, the corresponding validity verification rules are invoked to verify the multimodal interaction signals and filter out the valid trigger types.
3. The method according to claim 2, characterized in that, For the identified trigger types, the corresponding validity verification rules are invoked to verify the multimodal interaction signals and filter out valid trigger types, including: The verification threshold corresponding to the trigger type is retrieved. The verification threshold for the quick interaction type includes the action parameter threshold and the scene adaptation threshold, and the verification threshold for the deep interaction type includes the operation confirmation threshold. The characteristic parameters of the multimodal interaction signal are compared with the verification threshold to initially screen out candidate trigger signals that meet the parameter standards; The candidate trigger signals are subjected to scenario adaptation verification based on the current usage scenario of the user equipment. Trigger types that pass parameter comparison and scenario adaptation verification are determined to be valid trigger types.
4. The method according to claim 3, characterized in that, The step of performing scenario adaptation verification on the candidate trigger signals based on the current usage scenario of the user equipment includes: Real-time collection of multi-dimensional scene data from user devices, including at least environmental feature data, device status feature data, and user behavior feature data; The current usage scenario corresponding to the multi-dimensional scenario data is determined based on preset scenario classification rules. The current usage scenario includes any one of low threshold triggering scenario, high threshold triggering scenario, and prohibited triggering scenario. The candidate trigger signals are subjected to scenario adaptation verification based on the preset verification rules of the current usage scenario.
5. The method according to claim 1, characterized in that, Launching the corresponding target dialog window based on the valid trigger type includes: Based on the preset scene mapping relationship between trigger types and dialog windows, the target dialog window corresponding to the effective trigger type is determined; Retrieve the initial configuration parameters of the window corresponding to the target dialog window; Based on the initial window configuration parameters, the target dialog window is launched at a preset page position. When the target dialog window is a temporary dialog window, the preset page position is the top center or bottom pop-up area of the current device's first-level page, and the pop-up edge is reserved with a preset distance from the screen edge to avoid obscuring the core function entrance of the page. When the target dialog window is a regular dialog window, the preset page position is displayed as an independent new page in the current application level, with the original page return entrance retained at the top of the page, and the window page includes at least a history area, an input area, and a function area.
6. The method according to claim 1, characterized in that, The method of dynamically adjusting the display form of the target dialogue window based on the interactive content features generated during the dialogue includes: Real-time collection of interactive content during the dialogue process, including user speech-to-text content, AI reply text content, and multimedia interactive elements; Analyze the length, number of characters, layout requirements, and element types of the interactive content, and calculate the corresponding minimum display space; If the target dialog window is a temporary dialog window, the window height is dynamically adjusted based on the minimum display space so that the adjusted window height does not exceed the preset height percentage of the device screen height, and the interactive content is scrolled and displayed within the temporary dialog window after the window height is adjusted. If the target dialog window is a regular dialog window, the window height or width is adaptively adjusted based on the minimum display space so that the adjusted window height supports the complete display of interactive content.
7. The method according to claim 1, characterized in that, The preset switching conditions include at least one of the following: the adjusted window height reaches a preset height percentage of the device screen height, the length of the interactive content exceeds a preset number of characters, the user actively initiates a switching command, or multiple rounds of interaction are triggered. If the target dialog window is a temporary dialog window, then when the preset switching conditions are met, data is synchronously exchanged between the temporary dialog window and the regular dialog window, and a smooth switch is completed, including: Real-time monitoring of the height changes, content length, and user operation commands of the temporary dialog window to determine whether the preset switching conditions are met; When it is determined that the preset switching conditions are met, a window switching instruction is generated and the current interactive state of the temporary dialog window is frozen; Establish a data synchronization channel between the temporary dialog window and the regular dialog window; Based on the data synchronization channel, the temporary dialogue window is smoothly switched to the regular dialogue window according to the preset transition animation, and the interactive state is restored after the interactive data is synchronously loaded in the regular dialogue window.
8. An AI dialogue interaction device, characterized in that, include: The identification module is used to identify the valid trigger type corresponding to the multimodal trigger request in response to receiving a multimodal trigger request initiated by the user. The valid trigger type includes fast interaction trigger and deep interaction trigger. The startup module is used to launch the corresponding target dialog window according to the effective trigger type, wherein the quick interaction type trigger launches a temporary dialog window, and the deep interaction type trigger launches a regular dialog window; The adjustment module is used to dynamically adjust the display form of the target dialogue window based on the interactive content features generated during the dialogue process; The switching module is used to synchronously exchange data between the temporary dialog window and the regular dialog window and complete a smooth switch when the preset switching conditions are met if the target dialog window is a temporary dialog window.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
10. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.