Browser sidebar intelligent assistant interaction method and system

By fusing and reasoning from multiple data sources to generate a task recommendation queue, and by adjusting interaction options in conjunction with user sentiment data, the browser assistant achieves personalized and intelligent adaptation, thereby improving the accuracy of task recommendations and the interactive experience.

CN120949973APending Publication Date: 2025-11-14HEFEI D2S INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511463772.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, browser assistants lack the ability to combine user interaction behavior and emotions, resulting in a lack of targeted task recommendations and a low degree of matching between the interactive experience and the user's state and preferences.

Method used

By generating a task recommendation queue through multi-source data fusion and inference, and dynamically adjusting the display style of interactive options in combination with user sentiment data, the visualization template is accurately selected based on task results and sentiment data, realizing personalized and intelligent adaptation of the entire process from task recommendation to interactive feedback.

Benefits of technology

It significantly improves the accuracy and smoothness of interaction, makes task recommendations more aligned with user needs and real-time status, and closely revolves around user personality in the parsing, execution, and result rendering stages, thus solving the problem of low targeting and matching accuracy of traditional browser assistants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949973A_ABST
    Figure CN120949973A_ABST
Patent Text Reader

Abstract

The invention discloses a browser sidebar intelligent assistant interaction method and system, and relates to the technical field of man-machine interaction, and the method comprises the steps: obtaining webpage data according to webpage context information, obtaining interaction data from a user interaction record, and obtaining user emotion data through the interaction data; generating a task recommendation queue based on the webpage data, the interaction data and the emotion data by using a multi-dimensional reasoning method, and generating candidate interaction items according to the task recommendation queue and the user emotion data; receiving an instruction input by a user through the candidate interaction item, generating an analysis result in combination with the webpage data, executing a preset task corresponding to the analysis result, and obtaining task result data; and selecting a visualization template according to the task result data and the user emotion data to output a task result, and updating the user personalized configuration and the semantic memory library. The method is used for solving the problems that traditional browser assistant task recommendation lacks pertinence, and the matching degree of interactive experience and user states and preferences is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and more specifically, to a method and system for interacting with a browser sidebar intelligent assistant. Background Technology

[0002] With the rapid development of artificial intelligence and human-computer interaction technologies, browsers, as the primary entry point for users to access internet information, are gradually taking on a more intelligent and proactive service role. Especially against the backdrop of information overload and diversified tasks, traditional browser interaction methods based on static interfaces can no longer meet users' demands for efficient, personalized, and context-aware services. To improve user interaction efficiency and task completion experience during web browsing, sidebar intelligent assistants, as a browser enhancement component integrating semantic understanding and task management capabilities, are becoming an important development direction for intelligent browsers.

[0003] With the development of multimodal artificial intelligence, it has become possible to integrate information such as speech, text, and images for deep semantic understanding. By constructing a page representation model with structured and semantic features and combining it with user behavior data for modeling, more accurate intent recognition and task recommendation support can be provided for intelligent assistants. However, there is currently a lack of a unified method that can combine page structure, semantic information, user behavior, and user emotions in a browser environment to form a dynamic, scalable, and multi-dimensional integrated interactive system.

[0004] For example, the invention patent announcement CN119149060A discloses an injection method, apparatus, device, and computer-readable medium for an artificial intelligence assistant, relating to the field of artificial intelligence technology. The invention includes: performing page detection through a browser plugin included in the system to identify key feature information of the page, wherein the browser plugin is built based on a standardized API set; obtaining the user's context information in the current business based on the key feature information; injecting the artificial intelligence assistant into the current webpage according to the listening state of the browser plugin, and isolating the user interface components of the artificial intelligence assistant from the main information location of the current page based on an isolation container; obtaining user access requests, and responding to the user access requests by querying the context information of the current business and the user's historical access data through the artificial intelligence assistant.

[0005] For example, the invention patent announcement CN114995636A, concerning a multimodal interaction method and apparatus, relates to the field of computer technology. The method includes: receiving multimodal data, wherein the multimodal data includes voice data and video data; identifying the multimodal data to obtain user intent data and / or user posture data, wherein the user posture data includes user emotion data and user action data; determining a virtual character interaction strategy based on the user intent data and / or user posture data, wherein the virtual character interaction strategy includes a text interaction strategy and / or an action interaction strategy; acquiring a 3D rendering model of the virtual character; and generating an image of the virtual character containing the action interaction strategy using the 3D rendering model based on the virtual character interaction strategy, thereby driving the virtual character to perform multimodal interaction. This results in low latency throughout the interaction process and a better user experience.

[0006] The aforementioned publicly disclosed technical solutions have at least the following technical problems: CN119149060A proposes injecting AI assistants by recognizing page context information, but it does not fully consider user interaction behavior with web pages and real-time user emotions. CN114995636A proposes receiving multimodal data to determine interaction strategies, but it has not been implemented in specific scenarios, lacks connection with specific data, and relies too much on user action data, lacking initiative.

[0007] To address the above problems, this invention proposes a solution. Summary of the Invention

[0008] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a browser sidebar intelligent assistant interaction method and system. This method generates a task recommendation queue through multi-source data fusion and inference, dynamically adjusts the display style of interactive options based on user emotion data, and accurately selects visualization templates based on task results, user preferences, and emotion data. This achieves personalized and intelligent adaptation throughout the entire process from task recommendation to interactive feedback, effectively solving the problems of traditional browser assistant task recommendations lacking specificity and low matching degree between interactive experience and user status and preferences.

[0009] To achieve the above objectives, the present invention provides the following technical solution: A browser sidebar smart assistant interaction method includes the following steps: obtaining webpage data based on webpage context information, obtaining interaction data from user interaction records, and acquiring user sentiment data through the interaction data; using a multi-dimensional reasoning method, generating a task recommendation queue based on webpage data, interaction data, and sentiment data, and generating candidate interaction items based on the task recommendation queue and user sentiment data; receiving instructions input by the user through candidate interaction items, generating parsing results by combining webpage data, executing preset tasks corresponding to the parsing results, and obtaining task result data; selecting a visualization template based on the task result data and user sentiment data, outputting the task results, and updating the user's personalized configuration and semantic memory.

[0010] In a preferred embodiment, obtaining webpage context information specifically involves: obtaining webpage data based on the webpage context information, specifically: the webpage data includes page structure and semantic annotation data; the page structure is obtained by extracting page node hierarchy relationships, visibility status, and layout information to construct a page frame position element model; the semantic annotation data is obtained by identifying semantic value elements in the page, annotating text content and visual objects, and generating a semantic identifier matrix.

[0011] In a preferred embodiment, obtaining user emotion data through interaction data specifically involves: extracting high-frequency scene types from the interaction data and labeling typical emotion types in those scenes as baseline emotions; obtaining emotion classification results based on the interaction data using a multimodal emotion classification method; and correcting and optimizing the emotion classification results according to the baseline emotions to obtain user emotion data.

[0012] In a preferred embodiment, obtaining the emotion classification result based on interaction data using a multimodal emotion classification method specifically involves: the interaction data including text dialogues, voice commands, and interface operation events; converting text data into word feature vectors, extracting acoustic features from voice data, converting audio into audio feature vectors, converting operation events into numerical features, and generating operation feature vectors by combining semantic tags of web page elements; dividing emotion clusters using a clustering algorithm based on the interaction data feature vectors, labeling each emotion cluster with its corresponding emotion tag; calculating the similarity between its feature vector and the center of each emotion cluster based on real-time interaction data, determining the emotion cluster to which it belongs, and outputting the emotion classification result.

[0013] In a preferred embodiment, the step of correcting and optimizing the emotion classification results based on the baseline emotion to obtain user emotion data specifically involves: constructing a machine learning model to obtain the output emotion of a multimodal emotion classification model; allowing the machine learning model to learn the mapping relationship from the output emotion of the multimodal emotion classification model, the baseline emotion, and related interaction data to the real emotion; and using the trained machine learning model to correct and optimize the emotion classification results.

[0014] In a preferred embodiment, generating the task recommendation queue based on webpage data and interaction data specifically involves: generating an intent embedding vector based on the webpage data and interaction data; fusing the emotion data with the intent embedding vector to obtain a joint vector; calculating the similarity matrix between the joint vector and multimodal semantic tags in a pre-built task tag library to obtain a preliminary candidate task set; calculating weight factors based on the similarity matrix and historical interaction data, wherein the weight factors include scene association weight, frequency weight, and context weight; and sorting the preliminary candidate tasks using a weighted priority sorting algorithm based on the aforementioned weight factors to output the task recommendation queue.

[0015] In a preferred embodiment, generating candidate interaction items based on the task recommendation queue and user sentiment data specifically involves: dynamically generating multiple candidate interaction items based on the task recommendation queue, with each option encapsulated as an independent interaction card; monitoring the user input stream in real time and extracting input feature vectors, including input device type, screen resolution, and semantic complexity of the input content; and dynamically adjusting the card display style based on the input feature vectors and user sentiment data.

[0016] In a preferred embodiment, the step of receiving user input via candidate interaction items, generating parsing results by combining webpage data, executing a preset task corresponding to the parsing results, and obtaining task result data specifically involves: monitoring user input via candidate interaction items and converting it into structured text; merging the structured text with webpage data into an execution instruction set; performing difference matching between the execution instruction set and the semantic understanding module in the task tag library to generate parsing results; and matching the corresponding preset task based on the parsing results and calling the task execution module to complete the operation task.

[0017] In a preferred embodiment, the step of selecting a visualization template based on task result data and user sentiment data to output the task results and update the user's personalized configuration and semantic memory specifically involves: generating a preference weight matrix based on the current task result data and the user's historical preference configuration; calculating the priority scores of various display methods based on the preference weight matrix; selecting a target template from the rendering template library based on the priority scores and user sentiment data, and outputting the results to the template structure area; extracting task execution feedback data based on the user's response behavior and selection preferences to interactive content; and updating the user's personalized configuration and semantic memory by combining the task execution feedback data with the selected target template.

[0018] A browser sidebar intelligent assistant interaction system includes: a data acquisition and processing module, used to obtain webpage data based on webpage context information, obtain interaction data from user interaction records, and obtain user sentiment data through the interaction data; an intent prediction and interaction generation module, used to generate a task recommendation queue based on webpage data, interaction data, and sentiment data using a multi-dimensional reasoning method, and generate candidate interaction items based on the task recommendation queue and user sentiment data; a semantic parsing and execution module, used to receive instructions input by the user through candidate interaction items, generate parsing results by combining webpage data, execute preset tasks corresponding to the parsing results, and obtain task result data; and a result rendering and semantic memory update module, used to select a visualization template based on task result data and user sentiment data to output task results, and update user personalized configurations and semantic memory.

[0019] The technical effects and advantages of the browser sidebar smart assistant interaction method and system of the present invention are as follows: This invention generates a task recommendation queue through multi-source data fusion and inference, dynamically adjusts the display style of interactive options based on user emotion data, and accurately selects visualization templates based on task results and emotion data. This achieves personalized and intelligent adaptation throughout the entire process from task recommendation to interactive feedback. It builds a precise intent understanding foundation based on web page data, and integrates interactive and emotion data to make task recommendations more aligned with user needs and real-time status. Subsequent parsing, execution, and result rendering stages are also closely centered around user personality, significantly improving the accuracy and smoothness of interaction. This effectively solves the problems of traditional browser assistant task recommendations lacking specificity and low matching degree between interactive experience and user status and preferences. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a browser sidebar smart assistant interaction method according to the present invention.

[0021] Figure 2 This is a schematic diagram of the structure of a browser sidebar intelligent assistant interaction system according to the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0023] Example 1, Figure 1 The present invention provides a browser sidebar smart assistant interaction method, comprising the following steps: S1. Obtain webpage data based on webpage context information, obtain interaction data from user interaction records, and acquire user sentiment data through interaction data; S2 uses a multidimensional reasoning method to generate a task recommendation queue based on web page data, interaction data, and sentiment data, and generates candidate interaction items based on the task recommendation queue and user sentiment data. S3 receives the instructions input by the user through the candidate interaction items, combines the web page data to generate the parsing results, executes the preset tasks corresponding to the parsing results, and obtains the task result data; S4 selects a visualization template based on task result data and user sentiment data to output the task results and update the user's personalized configuration and semantic memory.

[0024] This embodiment generates a task recommendation queue through multi-source data fusion and inference, dynamically adjusts the display style of interactive options based on user sentiment data, and accurately selects visualization templates based on task results, user preferences, and sentiment data. This achieves personalized and intelligent adaptation throughout the entire process from task recommendation to interactive feedback, effectively solving the problems of traditional browser assistant task recommendations lacking specificity and low matching degree between interactive experience and user status and preferences.

[0025] S1: Obtain webpage data based on webpage context information, obtain interaction data from user interaction records, and acquire user sentiment data through interaction data.

[0026] In this embodiment, obtaining webpage data based on webpage context information specifically involves: The webpage context information includes page structure and semantic annotation data; Page structure is obtained by extracting page node hierarchy, visibility status, and layout information to construct a page frame position element model. Semantic annotation data is obtained by identifying semantic value elements on a page, annotating text content and visual objects, and generating a semantic tag matrix. Map webpage context information to webpage data.

[0027] The page frame position element model is represented by the following mathematical formula:

[0028] in, For a set of elements, The number of elements For the first A collection of information about each element. A unique identifier for an element. This represents the element's horizontal coordinates within the page viewport. This represents the element's vertical coordinates within the page viewport. The width of the element. The height of the element. The hierarchy of elements.

[0029] In this embodiment, mapping webpage context information to webpage data specifically involves: The page structure and semantic annotation data are jointly encoded to obtain information data units; Semantic fragments containing semantic features and spatial coordinates are generated through information data units; Semantic fragments are combined according to point-line-plane topology rules to generate web page data.

[0030] In this embodiment, the joint encoding of page structure and semantic annotation data to obtain information data units specifically involves: Parse the DOM tree of the webpage and extract spatial structure feature vectors; Perform semantic recognition on text and image elements in web pages and extract semantic feature vectors; Spatial structural feature vectors and semantic feature vectors are fused into information data units through attention weighting.

[0031] In this embodiment, the mathematical expression of the semantic segment is:

[0032] In the formula, For structural feature vectors, For semantic feature vectors, For behavioral characteristics, For location information, Features in the time dimension For the first A semantic fragment.

[0033] It should be noted that the structural feature vector (S) represents the element's hierarchy and parent-child relationship in the DOM, the semantic feature vector (M) represents the element's semantic type and the keywords and tags extracted by NLP, the behavioral feature (B) represents the element's interaction popularity, the location information (P) represents the element's position coordinates in the two-dimensional page, and the time dimension feature (T) represents dynamic dimensions such as the time of occurrence and duration of the behavior.

[0034] In this embodiment, the combination in the form of point-line-plane specifically refers to: Each semantic fragment is treated as a point, representing a semantic unit in a webpage. Each semantic fragment corresponds to an independent element or a region with specific semantics in the webpage. For semantic fragments that have a direct DOM hierarchy relationship, connection lines are constructed based on their hierarchy difference and connection type in the DOM tree. The order of semantic fragments involved in the user's interaction behavior on the webpage is recorded to form a user behavior path sequence. The connection lines and user behavior path sequences are weighted and merged to form logical and behavioral connections between nodes. Based on the behavioral clustering and spatial distribution of semantic fragments, they are aggregated into surfaces to form highly interactive areas or topic-focused areas.

[0035] It should be noted that the above data is compressed and combined using multi-bit binary or vector representation to form encoded segments with high information density and combinatorial expressive power.

[0036] In this embodiment, obtaining user emotion data through interaction data specifically includes: Extract high-frequency scene types from the interaction data and label the typical emotion types in the scene as the baseline emotion; Based on interactive data, emotion classification results are obtained through a multimodal emotion classification method; The emotion classification results are corrected and optimized based on the baseline emotion to obtain user emotion data.

[0037] In this embodiment, obtaining the emotion classification result based on interaction data using a multimodal emotion classification method specifically involves: The interactive data includes text dialogues, voice commands, and interface operation events; Text data is converted into word feature vectors, acoustic features of speech data are extracted, audio is converted into audio feature vectors, operation events are converted into numerical features, and operation feature vectors are generated by combining semantic tags of web page elements. Based on the feature vectors of the interaction data, a clustering algorithm is used to divide the emotion clusters and label the emotion tag corresponding to each emotion cluster. Based on real-time interactive data, the similarity between the feature vector and the center of each emotion cluster is calculated to determine the emotion cluster to which it belongs, and the emotion classification result is output.

[0038] In this embodiment, the step of correcting and optimizing the emotion classification results based on the baseline emotion to obtain user emotion data specifically involves: Build a machine learning model to obtain the output emotion of a multimodal emotion classification model; This allows machine learning models to learn the mapping relationship from the output emotion of a multimodal emotion classification model, the baseline emotion, and related interaction data to the real emotion. The trained machine learning model was used to refine and optimize the emotion classification results.

[0039] In this embodiment, the machine learning model is an LSTM model.

[0040] It should be noted that this embodiment achieves accurate semantic representation of web page content and intelligent perception of user emotions through refined coding modeling of web page structure and semantics, as well as multimodal sentiment analysis of user interaction data, providing a key basis for subsequent personalized adjustment of interaction.

[0041] S2 uses a multidimensional reasoning method to generate a task recommendation queue based on web page data, interaction data, and sentiment data, and generates candidate interaction items based on the task recommendation queue and user sentiment data.

[0042] In this embodiment, the method of using multidimensional reasoning to generate a task recommendation queue based on web page data, interaction data, and sentiment data specifically includes: Intent embedding vectors are generated based on web page data and interaction data, and sentiment data is fused with intent embedding vectors to obtain a joint vector. Calculate the similarity matrix between the joint vector and the multimodal semantic tags in the pre-built task tag library to obtain a preliminary set of candidate tasks; Weighting factors are calculated based on the similarity matrix and historical interaction data. These weighting factors include scene association weight, frequency weight, and context weight. Based on the aforementioned weighting factors, a weighted priority sorting algorithm is used to sort the initial candidate tasks and output a task recommendation queue.

[0043] It's important to note that webpage data can accurately pinpoint the task carrier and identify the specific object of task execution. It can precisely link these elements within the page, recommending more in-depth tasks and filtering tasks that contradict the current webpage context. Interaction data can capture user behavior patterns, prioritizing historically high-frequency tasks to reduce repetitive user operations. It also creates a unique behavioral profile for each user, making recommendations more aligned with individual preferences. Emotional data reflects the user's current emotional state. If the user is impatient, quick, one-step tasks are prioritized; if the user is calm, more in-depth, time-consuming tasks are recommended. Emotions influence a user's acceptance of interaction complexity. Incorporating emotional data into task recommendations can improve the user experience and alleviate negative emotions.

[0044] In this embodiment, the step of generating an intent embedding vector based on webpage data and interaction data, and fusing emotion data with the intent embedding vector to obtain a joint vector, specifically involves: Based on the core area of ​​the user's current interaction or the target range of the task, define the corresponding set of semantic fragments in the webpage; Extract page structure vectors, semantic annotation data vectors, and interface operation event vectors from semantic fragments within a defined range; The extracted vectors are weighted and fused using an attention network to obtain the intent embedding vector; The emotion data and the intention embedding vector are weighted and fused to form a joint vector.

[0045] In this embodiment, the specific formula for the similarity matrix between the joint vector and the multimodal semantic tags in the task tag library is as follows:

[0046] In the formula, For the intention vector and the first Semantic similarity between task tags For joint vectors, For task tags.

[0047] In this embodiment, the preliminary candidate task set is expressed mathematically as follows:

[0048] In the formula, This is a preliminary set of candidate tasks. Ranked Candidate task tags.

[0049] In this embodiment, three weighting factors are introduced for each candidate task label, and the specific formula is as follows:

[0050] In the formula, As a behavioral weighting factor, For users' past and tasks The success rate of interactive responses. For the task Frequency of appearance in the user's recent interaction history For tasks in the current page context semantic activation strength, These are the weighting coefficients for the three factors.

[0051] In this embodiment, the weighting coefficients of the three factors are dynamically adjusted as follows: Adjustments are made using a combination of base values ​​and dynamic compensation values. Based on the contextual weight, we detect whether there are time-limited or countdown elements on the webpage, detect whether the frequency of user operations increases suddenly, and determine the urgency of the scenario. Based on the frequency weight, we count the number of times the same task was executed in the past hour and the proportion of the task executed in the past week to determine the density of user behavior. Based on the context weight, we retrieve the historical satisfaction rating of the task, whether the user actively saved and reused the task results, and determine the validity of historical feedback. Thresholds are set for scenario urgency, user behavior density, and historical feedback validity. If the threshold is exceeded, the corresponding weighting coefficient is added to the base value and a compensation value is added; otherwise, the base value is used.

[0052] In this embodiment, the final task priority score is determined by both semantic similarity and behavioral weight, as shown in the following formula:

[0053] In the formula, For the first The priority score of each candidate task. For semantic similarity, This is the fusion coefficient between similarity and behavioral weight.

[0054] In this embodiment, the step of generating candidate interaction items based on the task recommendation queue and user sentiment data specifically includes: Based on the task recommendation queue, multiple candidate interaction items are dynamically generated, and each option is encapsulated as an independent interaction card. Real-time monitoring of user input stream, extraction of input feature vectors, including input device type, screen resolution, and semantic complexity of input content; The card display style is dynamically adjusted based on the input feature vector and user sentiment data.

[0055] In this embodiment, the step of dynamically adjusting the card display style based on the input feature vector and user emotional characteristics specifically includes rearranging the card grid layout based on screen resolution, folding or expanding the descriptive text according to semantic complexity, and adjusting the color saturation of the visual marker layer according to user emotional characteristics.

[0056] In this embodiment, the interactive card specifically includes: Task title, used to describe the task name; Task description, used to briefly explain the task content; Executable actions, direct manipulation of web page elements or functions corresponding to task tags; Highlight the target element, highlighting the relevant web page element area; Interactive buttons are buttons or other forms of interaction used to perform the task.

[0057] In this embodiment, the specific mathematical expression of the input feature vector is as follows:

[0058] in, For the input feature vector, The type of input device currently being used by the user. For the current screen resolution, The semantic complexity of the input content.

[0059] It should be noted that if the input feature vector contains If the value is less than a preset threshold, a vertical stacking method is used; if the value is greater than the preset threshold, a grid layout is used. If the value is below a preset threshold, a brief description will be displayed. If the value exceeds the preset threshold, the description will be expanded.

[0060] In this embodiment, the visual effect of the card is adjusted according to the user's emotional tendency. If the emotional tendency is positive, the card uses warm colors; if the emotional tendency is negative, the card color changes to cool colors to reduce visual stimulation.

[0061] It should be noted that this embodiment achieves accurate task recommendation driven by multi-source data through multi-dimensional reasoning method, and dynamically optimizes the display of interactive options in combination with user status, so that task recommendation is more in line with the scene, preferences and emotions, the interactive experience is more adapted to the user's real-time needs, and the interactive entry is more in line with the device characteristics and user status, which can effectively reduce operating costs and alleviate negative emotions.

[0062] S3 receives the instructions input by the user through candidate interactive items, combines the web page data to generate parsing results, executes the preset tasks corresponding to the parsing results, and obtains the task result data.

[0063] In this embodiment, receiving the user's instruction input through candidate interaction items, generating a parsing result by combining it with webpage data, executing the preset task corresponding to the parsing result, and obtaining the task result data specifically involves: Monitor user commands input through candidate interaction items and convert them into structured text; Merge structured text and web page data into an execution instruction set; The execution instruction set is matched with the semantic understanding module in the task tag library to generate parsing results; Based on the parsing results, the corresponding preset task is matched, the task execution module is called to complete the operation task and obtain the task result data.

[0064] In this embodiment, the specific formula for fusing structured text and webpage data into an execution instruction set is as follows:

[0065]

[0066] In the formula, For the merged execution instruction set, For the type of action the user wants to perform, The target object of the action. For additional parameter information, For the context of the current webpage, For structured text, This is webpage data.

[0067] In this embodiment, the step of performing a difference matching between the execution instruction set and each task tag in the task tag library, and calculating the semantic difference between the two, is specifically calculated using the following formula:

[0068] In the formula, To determine the difference between the instruction set and the task label, For the first in the task tag library A task label vector, and These are the vector norms of the execution instruction set and the task label, respectively.

[0069] In this embodiment, the parsing result includes user intent, action type, target object, and parameter information.

[0070] It should be noted that this embodiment achieves efficient parsing of user commands and accurate execution of tasks through instruction structure processing, multi-source data fusion, and precise difference matching, laying a reliable foundation for subsequent result output and providing clear guidance for subsequent task execution modules.

[0071] S4 selects a visualization template based on task result data and user sentiment data to output the task results and update the user's personalized configuration and semantic memory.

[0072] In this embodiment, the step of selecting a visualization template based on task result data and user sentiment data to output the task results and update the user's personalized configuration and semantic memory database specifically involves: Generate a preference weight matrix based on the current task result data and the user's historical preference configuration; Based on the preference weight matrix, the priority scores of various display methods are calculated; Select a target template from the rendering template library based on priority scores and user sentiment data, and output the result to the template structure area; Based on users' response behavior and selection preferences to interactive content, task execution feedback data is extracted; By combining task execution feedback data with the selected target template, the user's personalized configuration and semantic memory are updated.

[0073] In this embodiment, the specific formula for generating the preference weight matrix is ​​as follows:

[0074] In the formula, For the task Historical execution frequency For the task User satisfaction with the execution results For the task Feedback on the display method after execution. For preference weighting coefficients, Weights for user preferences.

[0075] In this embodiment, the preference weight coefficient is adjusted according to the task type, and different task types correspond to different preference weight coefficient values.

[0076] In this embodiment, the specific formula for calculating the priority score is as follows:

[0077] In the formula, Score based on priority. For display method Performance rating The total number of display methods, For the purpose of display method User preference weights.

[0078] In this embodiment, the step of selecting a target template from the rendering template library based on priority scores and user emotional characteristics specifically involves: Set mood-adaptive attribute tags for templates in the template library; By combining priority scores and user sentiment characteristics, a comprehensive matching weight is calculated for candidate templates in the template library that match the task type. The higher the weight, the higher the priority for selection. Based on the overall matching weight sorting, the template with the highest weight is selected as the target template. If there are templates with the same weight, the template with the higher historical usage frequency is selected.

[0079] In this embodiment, the comprehensive matching weight is specifically calculated using the following formula:

[0080] In the formula, To comprehensively match weights, These are the weighting coefficients. Score based on display method preference. The template's emotional fit score.

[0081] In this embodiment, the template emotion adaptation score is specifically set according to the emotion adaptation attribute tags set for the template library, with adaptation being 1 and non-adaptation being 0.

[0082] In this embodiment, the update of user personalized configuration is based on the following factors: Task type preference updates record user response behavior for different task types, such as whether users prefer quick query tasks or in-depth analysis tasks; Display preference updates record user preferences for different rendering templates, such as whether they prefer graphical or text summary displays, and whether they like interactive buttons.

[0083] In this embodiment, the specific formula for the preference update rule is as follows:

[0084] In the formula, Personalized configuration for users, Score the task type preference. Sensitivity to task response time The user satisfaction feedback score, These are the preference weighting coefficients.

[0085] In this embodiment, the specific formula for real-time updating of the user semantic memory is as follows:

[0086] In the formula, This is the current semantic memory vector. For learning rate, The changes in preferences resulting from the execution of this round of tasks. This is the updated semantic memory vector.

[0087] It should be noted that this embodiment achieves precise adaptation of task result output and continuous optimization of assistant interaction capabilities through template selection driven by both preferences and emotions and personalized memory updates. This allows the result rendering to match long-term preferences and real-time emotional states, and the assistant can continuously adapt to user habits as the number of interactions increases.

[0088] Example 2, Figure 2 The present invention discloses a browser sidebar intelligent assistant interaction system, comprising: The data acquisition and processing module is used to obtain web page data based on web page context information, obtain interaction data from user interaction records, and obtain user sentiment data through interaction data. The intent prediction and interaction generation module is used to generate a task recommendation queue based on web page data, interaction data and sentiment data using multi-dimensional reasoning methods, and to generate candidate interaction items based on the task recommendation queue and user sentiment data. The semantic parsing and execution module is used to receive instructions input by the user through candidate interaction items, combine them with web page data to generate parsing results, execute the preset tasks corresponding to the parsing results, and obtain task result data. The results rendering and semantic memory update module is used to select a visualization template based on task result data and user sentiment data to output task results and update user personalized configurations and semantic memory.

[0089] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0090] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0091] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0092] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0093] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0094] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A browser sidebar smart assistant interaction method, characterized in that, Includes the following steps: Web page data is obtained from web page context information, interaction data is obtained from user interaction records, and user sentiment data is obtained through interaction data. Using a multidimensional reasoning method, a task recommendation queue is generated based on web page data, interaction data, and sentiment data. Candidate interaction items are then generated based on the task recommendation queue and user sentiment data. It receives instructions from users through candidate interaction items, combines them with web page data to generate parsing results, executes the preset tasks corresponding to the parsing results, and obtains task result data. Based on task result data and user sentiment data, a visualization template is selected to output the task results, and the user's personalized configuration and semantic memory are updated.

2. The browser sidebar smart assistant interaction method according to claim 1, characterized in that, The process of obtaining webpage data based on webpage context information specifically includes: The webpage context information includes page structure and semantic annotation data; Page structure is obtained by extracting page node hierarchy, visibility status, and layout information to construct a page frame position element model. Semantic annotation data is obtained by identifying semantic value elements on a page, annotating text content and visual objects, and generating a semantic tag matrix. Map webpage context information to webpage data.

3. The browser sidebar smart assistant interaction method according to claim 2, characterized in that, The acquisition of user sentiment data through interactive data specifically includes: Extract high-frequency scene types from the interaction data and label the typical emotion types in the scene as the baseline emotion; Based on interactive data, emotion classification results are obtained through a multimodal emotion classification method; The emotion classification results are corrected and optimized based on the baseline emotion to obtain user emotion data.

4. The browser sidebar smart assistant interaction method according to claim 3, characterized in that, The emotion classification result obtained based on interactive data and through a multimodal emotion classification method is as follows: The interactive data includes text dialogues, voice commands, and interface operation events; Text data is converted into word feature vectors, acoustic features of speech data are extracted, audio is converted into audio feature vectors, operation events are converted into numerical features, and operation feature vectors are generated by combining semantic tags of web page elements. Based on the feature vectors of the interaction data, a clustering algorithm is used to divide the emotion clusters and label the emotion tag corresponding to each emotion cluster. Based on real-time interactive data, the similarity between the feature vector and the center of each emotion cluster is calculated to determine the emotion cluster to which it belongs, and the emotion classification result is output.

5. The browser sidebar smart assistant interaction method according to claim 4, characterized in that, The process of correcting and optimizing the emotion classification results based on the baseline emotion to obtain user emotion data is as follows: Build a machine learning model to obtain the output emotion of a multimodal emotion classification model; This allows machine learning models to learn the mapping relationship from the output emotion of a multimodal emotion classification model, the baseline emotion, and related interaction data to the real emotion. The trained machine learning model was used to refine and optimize the emotion classification results.

6. The browser sidebar smart assistant interaction method according to claim 5, characterized in that, The method of using multidimensional reasoning to generate a task recommendation queue based on web page data, interaction data, and sentiment data specifically involves: Intent embedding vectors are generated based on web page data and interaction data, and sentiment data is fused with intent embedding vectors to obtain a joint vector. Calculate the similarity matrix between the joint vector and the multimodal semantic tags in the pre-built task tag library to obtain a preliminary set of candidate tasks; Weighting factors are calculated based on the similarity matrix and historical interaction data. These weighting factors include scene association weight, frequency weight, and context weight. Based on the aforementioned weighting factors, a weighted priority sorting algorithm is used to sort the initial candidate tasks and output a task recommendation queue.

7. The browser sidebar smart assistant interaction method according to claim 6, characterized in that, The process of generating candidate interaction items based on the task recommendation queue and user sentiment data specifically involves: Based on the task recommendation queue, multiple candidate interaction items are dynamically generated, and each option is encapsulated as an independent interaction card. Real-time monitoring of user input stream, extraction of input feature vectors, including input device type, screen resolution, and semantic complexity of input content; The card display style is dynamically adjusted based on the input feature vector and user sentiment data.

8. The browser sidebar smart assistant interaction method according to claim 7, characterized in that, The process of receiving user input via candidate interactive items, generating parsing results by combining webpage data, executing preset tasks corresponding to the parsing results, and obtaining task result data specifically involves: Monitor user commands input through candidate interaction items and convert them into structured text; Merge structured text and web page data into an execution instruction set; The execution instruction set is matched with the semantic understanding module in the task tag library to generate parsing results; Based on the parsing results, the corresponding preset task is matched, the task execution module is called to complete the operation task and obtain the task result data.

9. The browser sidebar smart assistant interaction method according to claim 8, characterized in that, The process of selecting a visualization template based on task result data and user sentiment data to output task results, and updating user personalized configurations and semantic memory database, specifically involves: Generate a preference weight matrix based on the current task result data and the user's historical preference configuration; Based on the preference weight matrix, the priority scores of various display methods are calculated; Select a target template from the rendering template library based on priority scores and user sentiment data, and output the result to the template structure area; Based on users' response behavior and selection preferences to interactive content, task execution feedback data is extracted; By combining task execution feedback data with the selected target template, the user's personalized configuration and semantic memory are updated.

10. A system using a browser sidebar smart assistant interaction method as described in any one of claims 1-9, characterized in that, include: The data acquisition and processing module is used to obtain web page data based on web page context information, obtain interaction data from user interaction records, and obtain user sentiment data through interaction data. The intent prediction and interaction generation module is used to generate a task recommendation queue based on web page data, interaction data and sentiment data using multi-dimensional reasoning methods, and to generate candidate interaction items based on the task recommendation queue and user sentiment data. The semantic parsing and execution module is used to receive instructions input by the user through candidate interaction items, combine them with web page data to generate parsing results, execute the preset tasks corresponding to the parsing results, and obtain task result data. The results rendering and semantic memory update module is used to select a visualization template based on task result data and user sentiment data to output task results and update user personalized configurations and semantic memory.

Citation Information

Patent Citations

  • Multi-modal interaction method and device

    CN114995636A

  • Artificial intelligence assistant injection method, device and equipment and computer readable medium

    CN119149060A

Cited By

  • Multi-model intelligent interaction method and system based on browser extension and computing equipment

    CN121657897A