Method and system for automatic optimization of software operation prompts based on context understanding
By constructing multi-dimensional context feature data and using multi-modal feature fusion network and bidirectional long and short-term memory network, predicting user operation intentions and optimizing prompt generation strategies, the problem that prompt content cannot be dynamically adjusted in the existing technology is solved, and highly targeted personalized prompts are achieved.
Patent Information
- Application Number
- CN202510144049.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Existing software operation prompt technology cannot dynamically adjust the prompt content based on user usage habits and operation characteristics, resulting in insufficient targeted and personalized prompt information and lack of effective prompt optimization mechanism.
By collecting and constructing multi-dimensional context feature data for user software operations, using multi-modal feature fusion network and bidirectional long and short-term memory network for feature encoding and timing-dependent modeling, the dynamic evolution law of user operation behavior is captured. Based on the multi-head attention mechanism, the fused context semantic representation vector is generated, the user's operation intention and operation probability distribution is predicted, and the prompt generation strategy is optimized through reinforcement learning methods.
It realizes accurate understanding and prediction of user operation intentions, improves the pertinence and accuracy of software operation prompts, and the generated prompt information is more personalized and practical. The system can adaptively learn user preferences and continuously improve prompt effects.
Smart Images

Figure CN119598205B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, and in particular to a method and system for automatically optimizing software operation prompts based on context understanding. Background Art
[0002] With the increasing complexity and functional diversification of software systems, in order to help users better use software functions, various software systems generally use operation prompt functions to assist users. Traditional software operation prompts usually use preset fixed prompt templates to trigger corresponding prompt information when users perform specific operations or enter specific interfaces. This prompt method can provide users with basic operation guidance and help users understand the basic functions and operation methods of the software. At present, some software systems have also introduced rule-based intelligent prompt mechanisms to achieve more targeted operation prompts by setting trigger conditions and prompt rules.
[0003] However, the existing software operation prompt technology still has the following shortcomings:
[0004] Existing operation prompt methods often adopt static prompt strategies, which cannot dynamically adjust the prompt content according to the user's usage habits and operation characteristics, resulting in insufficient pertinence and personalization of the prompt information. The prompt content is often a pre-set fixed text, lacking an in-depth understanding of the user's current operation context, and unable to accurately grasp the user's actual needs and operation intentions.
[0005] Traditional prompt systems usually only consider single-dimensional trigger conditions, such as interface status or operation sequence, and ignore the multi-dimensional contextual information behind user operation behaviors, including historical operation trajectories, environmental factors, etc. This makes it impossible for the prompt system to fully understand the user's operation scenario, reducing the accuracy and practicality of the prompt.
[0006] Existing technologies lack an effective prompt optimization mechanism and are unable to continuously improve and optimize prompt strategies based on user feedback. Most prompt systems use fixed prompt rules and lack adaptive learning capabilities, making it difficult to dynamically adjust as user usage habits change, resulting in difficulty in continuously improving prompt effects. Summary of the invention
[0007] The embodiments of the present invention provide a method and system for automatically optimizing software operation prompts based on context understanding, which can solve the problems in the prior art.
[0008] A first aspect of an embodiment of the present invention provides a method for automatically optimizing software operation prompts based on context understanding, comprising:
[0009] Collect and construct a multi-dimensional context feature data set of user software operations, wherein the multi-dimensional context feature data set includes at least one of historical operation sequence data of the user within a preset time window, interface state data of the current software, real-time user interaction behavior data, and software operating environment data; perform data standardization and time sequence alignment processing on the multi-dimensional context feature data set to generate pre-processed context feature data;
[0010] The preprocessed context feature data is input into a multimodal feature fusion network, and feature encoding is performed on historical operation sequence data, interface status data, user real-time interaction behavior data, and software operating environment data, respectively, to generate a multimodal feature vector; a bidirectional long short-term memory network is used to model the time-series dependency of the multimodal feature vector, to capture the dynamic evolution law of user operation behavior, and output the time-series context feature; based on a multi-head attention mechanism, the importance of different dimensions of the time-series context feature is weighted to generate a fused context semantic representation vector; the context semantic representation vector is input into an intention classification network to predict the user's operation intention and operation probability distribution;
[0011] Based on the user's operation intention, similarity matching is performed in the prompt strategy library, and the prompt template with the highest matching degree is selected as the basic template; the operation probability distribution is used as the importance weight of the prompt content, and the prompt elements in the basic template are prioritized; the prompt generation strategy is continuously optimized using a reinforcement learning method, and the user's interactive feedback data on historical prompt information is used as a reward signal. The dynamic adjustment rules of the prompt content are learned through a strategy network; the basic template is rewritten and supplemented according to the dynamic adjustment rules to generate personalized prompt information that conforms to the current context.
[0012] The historical operation sequence data records the access order of function modules and the timestamp of operation events; the interface status data of the current software includes the hierarchical structure information of the interface function modules and the spatial distribution information of the interactive controls; the real-time user interactive behavior data includes the coordinate sequence of the mouse movement trajectory, the operation dwell time statistics and the control triggering frequency; the software operating environment data includes the software version identification information, system configuration parameters and hardware equipment information.
[0013] The multimodal feature vector is modeled for temporal dependency using a bidirectional long short-term memory network to capture the dynamic evolution of user operation behavior and output temporal context features; the different dimensions of the temporal context features are weighted based on a multi-head attention mechanism to generate a fused context semantic representation vector; the context semantic representation vector is input into an intention classification network to predict user operation intention and operation probability distribution, including:
[0014] Inputting the multimodal feature vector into a bidirectional long short-term memory network, the bidirectional long short-term memory network includes a forward propagation layer and a backward propagation layer, and performing time series modeling on the multimodal feature vector through the forward propagation layer and the backward propagation layer, respectively, to obtain high-level time series features, middle-level time series features and fine-grained time series features, wherein the high-level time series features are used to characterize function module selection information, the middle-level time series features are used to characterize interface navigation information, and the fine-grained time series features are used to characterize control operation information;
[0015] Generate a query matrix, a key matrix and a value matrix based on the high-level temporal features, calculate the similarity between the query matrix and the key matrix to obtain a high-level attention weight, and weight the high-level attention weight and the value matrix to obtain a high-level context vector; splice the high-level context vector with the middle-level temporal features to generate a conditional query matrix, calculate the middle-level attention weight based on the conditional query matrix and obtain the middle-level context vector; splice the middle-level context vector with the fine-grained temporal features to generate a fine-grained conditional query matrix, calculate the fine-grained attention weight based on the fine-grained conditional query matrix and obtain the fine-grained context vector;
[0016] Input the high-level context vector into the first fully connected layer of the intent classification network to obtain a high-level intent probability distribution; based on the combined features of the high-level intent probability distribution and the middle-level context vector, obtain the middle-level intent probability distribution through the second fully connected layer; based on the combined features of the high-level intent probability distribution, the middle-level intent probability distribution and the fine-grained context vector, obtain the fine-grained intent probability distribution through the third fully connected layer;
[0017] Calculate the cross entropy loss of the high-level intent probability distribution, the mid-level intent probability distribution and the fine-grained intent probability distribution relative to the true label, and calculate the relative entropy between adjacent level intent probability distributions to obtain the hierarchical consistency constraint term; perform a weighted combination of the cross entropy loss and the hierarchical consistency constraint term to construct a joint loss function, and optimize the parameters of the bidirectional long short-term memory network, the first fully connected layer, the second fully connected layer and the third fully connected layer through the gradient descent algorithm.
[0018] Based on the user operation intention, similarity matching is performed in the prompt strategy library, and the prompt template with the highest matching degree is selected as the basic template; the operation probability distribution is used as the importance weight of the prompt content, and the prompt elements in the basic template are prioritized, including:
[0019] Constructing a multi-level prompt strategy library, wherein the multi-level prompt strategy library includes a plurality of strategy templates, each of which includes a template body text, a prompt element set, and an element dependency matrix, wherein the element dependency matrix is used to characterize the association relationship between each prompt element in the prompt element set;
[0020] Using a first semantic encoder to encode the user operation intention to obtain an intention semantic vector, using a second semantic encoder to encode the template body text of each of the policy templates in the multi-level prompt policy library to obtain a template semantic vector; calculating the cosine similarity between the intention semantic vector and the template semantic vector to obtain a first similarity component, calculating the correlation between the intention semantic vector and each prompt element in the prompt element set to obtain a second similarity component, performing a weighted combination of the first similarity component and the second similarity component to obtain a comprehensive similarity, selecting the best matching policy template from the multi-level prompt policy library based on the comprehensive similarity, and selecting the prompt template with the highest matching degree as the basic template;
[0021] Constructing an intention factor mapping matrix, performing matrix multiplication operation on the operation probability distribution and the intention factor mapping matrix to obtain an initial factor weight vector; performing weighted combination on the initial factor weight vector and the factor dependency matrix to obtain a fused factor weight vector, wherein the fused factor weight vector considers both intention relevance and factor dependency;
[0022] Based on the fusion element weight vector, the prompt element set in the optimal matching strategy template is sorted by importance to obtain an ordered prompt element sequence; according to the ordered prompt element sequence and the fusion element weight vector, the template body text of the optimal matching strategy template is reconstructed to generate importance-aware adaptive prompt content.
[0023] The prompt generation strategy is continuously optimized by using the reinforcement learning method, and the interactive feedback data of the user on the historical prompt information is used as the reward signal. The dynamic adjustment rules of the prompt content are learned through the strategy network; the basic template is rewritten and supplemented according to the dynamic adjustment rules to generate personalized prompt information that conforms to the current context, including:
[0024] Construct a multi-dimensional reward function for user feedback, wherein the multi-dimensional reward function includes a click response reward, a stay time reward, a task completion reward, and a satisfaction score reward, and calculate the user's interactive feedback score for the historical prompt information based on the multi-dimensional reward function;
[0025] Acquire current context features, historical interaction sequence codes, prompt template features, and user portrait features, and combine the current context features, the historical interaction sequence codes, the prompt template features, and the user portrait features to generate a state vector; construct an action space, the action space including modification operations, append operations, deletion operations, and reordering operations, and each operation includes an operation type parameter, a location information parameter, and a content feature parameter;
[0026] Constructing a policy network and a value network, wherein the policy network is used to predict the execution probability distribution of each operation according to the state vector, and the value network is used to evaluate the value score of the state vector; calculating the advantage function value based on the interactive feedback score and the value score, performing policy gradient update on the parameters of the policy network according to the advantage function value, and performing value estimation update on the parameters of the value network according to the temporal difference error;
[0027] Constructing a prompt adjustment rule set based on the updated policy network, wherein each rule in the prompt adjustment rule set includes a trigger condition, an adjustment action, and a rule priority; calculating an execution probability of each adjustment action according to the state vector, and multiplying the execution probability by a user personalized weight to obtain a comprehensive execution priority;
[0028] Based on the comprehensive execution priority, an optimal adjustment action is selected from the action space, and the prompt template is dynamically adjusted according to the optimal adjustment action to generate personalized prompt content; the user's interactive feedback data on the personalized prompt content is input into the multi-dimensional reward function, and the parameters of the policy network and the value network are continuously optimized to realize online learning of the prompt strategy.
[0029] A prompt adjustment rule set is constructed based on the updated policy network, wherein each rule in the prompt adjustment rule set includes a trigger condition, an adjustment action, and a rule priority; an execution probability of each adjustment action is calculated according to the state vector, and a comprehensive execution priority is obtained by multiplying the execution probability by the user personalized weight, including:
[0030] Constructing a rule representation model, wherein the rule representation model includes a rule trigger condition set, an adjustment action description, a rule priority, and a rule credibility, wherein each trigger condition in the rule trigger condition set includes a feature attribute, a comparison operator, and a threshold parameter, and the rule trigger condition set is used to filter the state vector;
[0031] The state vector is processed by using a policy network to generate candidate actions and their corresponding action probabilities, the gradient information of the state vector relative to each candidate action is calculated, and the state-action pair whose action probability is greater than a preset probability threshold and whose gradient significance is greater than a preset gradient threshold is used as an initial adjustment rule;
[0032] The initial adjustment rule is evaluated, rule confidence is calculated based on historical samples that meet the rule triggering condition set, rule support is calculated based on the historical triggering times of the initial adjustment rule, rule novelty is calculated based on the difference between the initial adjustment rule and the existing rule set, and rule confidence, rule support and rule novelty are weightedly combined to obtain rule priority;
[0033] Acquire user interaction data, extract multi-dimensional user features from the user interaction data to construct a user preference vector, map the candidate actions to a feature space using an action mapping function, map the user preference vector to the same feature space, and calculate a personalized weight based on the matching degree of the two mapping results;
[0034] Multiplying the action probability by the personalized weight and introducing a decay factor based on the complexity of the action to obtain a comprehensive execution priority of each of the candidate actions, wherein the complexity of the action is determined according to the number of operation steps involved in the action;
[0035] The rule set is dynamically maintained based on the comprehensive execution priority and the rule priority, a final priority score of each rule in the rule set is calculated, rules whose final priority scores are higher than a priority threshold and whose usage frequencies are higher than a usage frequency threshold are retained in the rule set, and a personalized prompt adjustment strategy is generated based on the rule set.
[0036] A second aspect of an embodiment of the present invention provides a system for automatically optimizing software operation prompts based on context understanding, including:
[0037] The first unit is used to collect and construct a multi-dimensional context feature data set of user software operations, wherein the multi-dimensional context feature data set includes at least one of historical operation sequence data of the user within a preset time window, interface state data of the current software, real-time user interaction behavior data, and software operating environment data; and perform data standardization and time sequence alignment processing on the multi-dimensional context feature data set to generate pre-processed context feature data;
[0038] The second unit is used to input the preprocessed context feature data into a multimodal feature fusion network, perform feature encoding on historical operation sequence data, interface state data, user real-time interaction behavior data and software operating environment data, and generate a multimodal feature vector; use a bidirectional long short-term memory network to perform temporal dependency modeling on the multimodal feature vector, capture the dynamic evolution law of user operation behavior, and output temporal context features; based on a multi-head attention mechanism, the importance of different dimensions of the temporal context features is weighted to generate a fused context semantic representation vector; the context semantic representation vector is input into an intention classification network to predict user operation intentions and operation probability distribution;
[0039] The third unit is used to perform similarity matching in the prompt strategy library based on the user operation intention, and select the prompt template with the highest matching degree as the basic template; use the operation probability distribution as the importance weight of the prompt content, and prioritize the prompt elements in the basic template; use the reinforcement learning method to continuously optimize the prompt generation strategy, use the user's interactive feedback data on historical prompt information as a reward signal, and learn the dynamic adjustment rules of the prompt content through the strategy network; rewrite and supplement the basic template according to the dynamic adjustment rules to generate personalized prompt information that conforms to the current context.
[0040] A third aspect of the embodiments of the present invention
[0041] An electronic device is provided, comprising:
[0042] processor;
[0043] a memory for storing processor-executable instructions;
[0044] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0045] According to a fourth aspect of the embodiments of the present invention,
[0046] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0047] The present invention has the following beneficial effects through the software operation prompt automatic optimization method based on context understanding:
[0048] By constructing a multi-dimensional contextual feature dataset and preprocessing it, combined with a multimodal feature fusion network and a bidirectional long short-term memory network, it is possible to comprehensively capture the user's operating behavior patterns and usage habits, accurately understand and predict the user's operating intentions, and improve the pertinence and accuracy of software operation prompts.
[0049] A multi-head attention mechanism is used to weight the importance of temporal context features, and the prompt elements are prioritized based on the operation probability distribution, so that the generated prompt information can highlight the key content, avoid redundant information from interfering with users, and improve the practicality and readability of the prompt information.
[0050] The prompt generation strategy is continuously optimized through reinforcement learning methods, and user feedback is used as a reward signal to dynamically adjust the prompt content, so that the system can adaptively learn user preferences and continuously improve the prompt effect, realizing personalized customization and dynamic optimization of software operation prompts, and significantly improving the user's software usage experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A flowchart of a method for automatically optimizing software operation prompts based on context understanding according to an embodiment of the present invention;
[0052] Figure 2 The present invention is a schematic diagram of the structure of a system for automatically optimizing software operation prompts based on context understanding according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0054] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0055] Figure 1 FIG. 1 is a flow chart of a method for automatically optimizing software operation prompts based on context understanding according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0056] S101. Collect and construct a multi-dimensional context feature data set of user software operations, wherein the multi-dimensional context feature data set includes at least one of the user's historical operation sequence data within a preset time window, the current software interface state data, the user's real-time interactive behavior data, and the software running environment data; perform data standardization and time sequence alignment processing on the multi-dimensional context feature data set to generate pre-processed context feature data;
[0057] S102. Input the preprocessed context feature data into the multimodal feature fusion network, perform feature encoding on the historical operation sequence data, interface state data, user real-time interaction behavior data and software running environment data respectively, and generate a multimodal feature vector; use a bidirectional long short-term memory network to perform temporal dependency modeling on the multimodal feature vector, capture the dynamic evolution law of user operation behavior, and output temporal context features; based on the multi-head attention mechanism, weight the importance of different dimensions of the temporal context features to generate a fused context semantic representation vector; input the context semantic representation vector into the intention classification network to predict the user's operation intention and operation probability distribution;
[0058] S103. Perform similarity matching in the prompt strategy library based on the user's operation intention, and select the prompt template with the highest matching degree as the basic template; use the operation probability distribution as the importance weight of the prompt content, and prioritize the prompt elements in the basic template; use the reinforcement learning method to continuously optimize the prompt generation strategy, use the user's interactive feedback data on historical prompt information as a reward signal, and learn the dynamic adjustment rules of the prompt content through the strategy network; rewrite and supplement the basic template according to the dynamic adjustment rules to generate personalized prompt information that conforms to the current context.
[0059] In an optional implementation, the historical operation sequence data records the access order of function modules and the timestamp of operation events; the interface status data of the current software includes the hierarchical structure information of the interface function modules and the spatial distribution information of the interactive controls; the real-time user interactive behavior data includes the coordinate sequence of the mouse movement trajectory, the operation dwell time statistics and the control triggering frequency; the software operating environment data includes the software version identification information, system configuration parameters and hardware equipment information.
[0060] The specific implementation of the historical operation sequence data recording the access sequence of functional modules and the timestamp of operation events adopts the event monitoring mechanism. By implanting a listener at the entrance of each functional module of the software, the module identifier, entry timestamp and exit timestamp when the user accesses are recorded. These data are stored in JSON format, including fields such as module ID, access start time, and end time. For example, the record example of a user accessing the system settings module: {"moduleId": "systemSettings", "enterTime": "2024-03-21 10:30:15", "exitTime": "2024-03-21 10:35:20"}.
[0061] The collection of the current software interface status data is achieved through DOM tree parsing. First, obtain the DOM structure of the entire interface, parse the parent-child relationship between nodes at each level, and build a hierarchical tree of interface function modules. At the same time, record the coordinate position, size and other spatial attributes of each interactive control. Example of interface status data: {"layout": {"header": {"height": 60, "children": ["logo", "menu"]}, "content": {"position": {"x": 0,"y": 60}, "width": 1200}}}.
[0062] The data collection of real-time user interaction behavior adopts event capture. By monitoring mouse movement events, the mouse coordinates are sampled every 50 milliseconds to form a trajectory coordinate sequence. For control operations, the user's dwell time on each control is recorded, and the number of triggers of each control is counted. Interaction data example: {"mouseTrack": [[100,200,0],[150,220,50],[180,225,100]], "dwellTime": {"button1": 2000, "input1": 5000}, "triggerCount": {"button1": 5, "button2": 2}}.
[0063] The software running environment data is obtained through the system API interface. Including reading the software version number, operating system type and version, CPU model, memory capacity, graphics card information, etc. Environment data example: {"softwareVersion": "2.1.0", "osInfo": {"type": "Windows", "version": "10"}, "hardware": {"cpu": "Inteli7", "memory": "16GB"}}.
[0064] All collected data is integrated and stored through a unified data processing module. First, the raw data is cleaned to remove outliers and redundant information. The processed data is then structured and stored in a predefined format for easy subsequent analysis and use. Data storage uses a distributed database to support fast reading and writing and historical data query.
[0065] This application can achieve:
[0066] By comprehensively recording user operation history and software operation status, detailed data support is provided for system optimization and problem diagnosis, improving software maintenance efficiency and service quality. Based on accurate interface status and user behavior data, it can deeply analyze user usage habits and operation patterns, provide reliable decision-making basis for product improvement, and enhance user experience. The use of distributed storage and real-time data processing mechanisms ensures efficient collection and processing of large-scale data, and the system runs stably and reliably with good scalability.
[0067] In an optional implementation, a bidirectional long short-term memory network is used to model the temporal dependency of the multimodal feature vector, capture the dynamic evolution law of the user's operation behavior, and output the temporal context features; based on a multi-head attention mechanism, the importance of different dimensions of the temporal context features is weighted to generate a fused context semantic representation vector; the context semantic representation vector is input into an intention classification network, and the prediction of the user's operation intention and operation probability distribution includes:
[0068] Inputting the multimodal feature vector into a bidirectional long short-term memory network, the bidirectional long short-term memory network includes a forward propagation layer and a backward propagation layer, and performing time series modeling on the multimodal feature vector through the forward propagation layer and the backward propagation layer, respectively, to obtain high-level time series features, middle-level time series features and fine-grained time series features, wherein the high-level time series features are used to characterize function module selection information, the middle-level time series features are used to characterize interface navigation information, and the fine-grained time series features are used to characterize control operation information;
[0069] Generate a query matrix, a key matrix and a value matrix based on the high-level temporal features, calculate the similarity between the query matrix and the key matrix to obtain a high-level attention weight, and weight the high-level attention weight and the value matrix to obtain a high-level context vector; splice the high-level context vector with the middle-level temporal features to generate a conditional query matrix, calculate the middle-level attention weight based on the conditional query matrix and obtain the middle-level context vector; splice the middle-level context vector with the fine-grained temporal features to generate a fine-grained conditional query matrix, calculate the fine-grained attention weight based on the fine-grained conditional query matrix and obtain the fine-grained context vector;
[0070] Input the high-level context vector into the first fully connected layer of the intent classification network to obtain a high-level intent probability distribution; based on the combined features of the high-level intent probability distribution and the middle-level context vector, obtain the middle-level intent probability distribution through the second fully connected layer; based on the combined features of the high-level intent probability distribution, the middle-level intent probability distribution and the fine-grained context vector, obtain the fine-grained intent probability distribution through the third fully connected layer;
[0071] Calculate the cross entropy loss of the high-level intent probability distribution, the mid-level intent probability distribution and the fine-grained intent probability distribution relative to the true label, and calculate the relative entropy between adjacent level intent probability distributions to obtain the hierarchical consistency constraint term; perform a weighted combination of the cross entropy loss and the hierarchical consistency constraint term to construct a joint loss function, and optimize the parameters of the bidirectional long short-term memory network, the first fully connected layer, the second fully connected layer and the third fully connected layer through the gradient descent algorithm.
[0072] First, the user's multimodal feature vector is preprocessed. The multimodal feature vector contains the visual features, text features, and behavioral features of the user interface operation. The visual features are extracted through the convolutional neural network to extract the image features of the interface screenshot, the text features are encoded through the word embedding model, and the behavioral features include numerical features such as click coordinates and operation time intervals. These features are normalized and spliced into a unified feature vector.
[0073] Then the preprocessed multimodal feature vector is input into the bidirectional long short-term memory network for time series modeling. The bidirectional long short-term memory network contains two propagation layers, forward and backward, each of which contains multiple long short-term memory units. Taking the time window of 10 time steps as an example, the forward propagation layer processes the feature sequence from left to right, and the backward propagation layer processes the feature sequence from right to left. The hidden states in the two directions capture temporal dependencies of different granularities respectively. Specifically, the shallow layer of the network encodes high-level semantic information such as function module selection, the middle layer encodes interface navigation information, and the deep layer encodes fine-grained control operation information.
[0074] Then, based on the multi-head attention mechanism, the importance of temporal features at different levels is weighted. Taking 8 attention heads as an example, the high-level temporal features are first linearly transformed to obtain the query matrix, key matrix and value matrix. The attention weight is obtained by calculating the similarity between the query matrix and the key matrix, and the high-level context vector is obtained by multiplying the weight with the value matrix. Similarly, after the high-level context vector is concatenated with the middle-level temporal features, the middle-level attention weight and context vector are calculated through conditional query. Finally, the middle-level context vector is concatenated with the fine-grained temporal features to obtain the fine-grained context vector.
[0075] Finally, a hierarchical intent classification network is used to predict user operation intentions. The network consists of three fully connected layers, which predict intentions of different granularities. The first layer predicts the function module selection intention based on the high-level context vector, the second layer combines the high-level intent probability and the middle-level context vector to predict the interface navigation intention, and the third layer integrates the multi-layer intent probability and the fine-grained context vector to predict the specific control operation intention. The cross entropy loss and hierarchical consistency constraint are used to construct the joint loss function, and the model parameters are optimized by the stochastic gradient descent algorithm.
[0076] For example, when a user performs continuous operations in the intelligent human-computer interaction system, the user's multimodal interaction data is first obtained, including features such as touch position coordinate sequence, interface state information, and operation control type, and these original features are converted into multimodal feature vectors of fixed dimensions through a feature extraction module. Take the operation sequence of a user in a smart home control application as an example, which includes continuous actions such as opening the application, clicking on the device list, selecting the air-conditioning device, and adjusting the temperature.
[0077] A bidirectional long short-term memory network is used to model the time series of these multimodal features. The network contains two propagation layers, forward and backward, each containing multiple memory units. Taking a 32-dimensional input feature vector as an example, the forward propagation layer processes the feature sequence from left to right to capture the contextual dependencies before the current moment; the backward propagation layer processes from right to left to capture the contextual information afterwards. The outputs of the two directions are mapped to a 256-dimensional hidden space, respectively, to obtain three levels of time series features:
[0078] The high-level temporal feature dimension is 256, which is used to characterize the user's overall function selection intention, such as "adjust the air conditioner temperature"; the middle-level temporal feature dimension is 128, which characterizes the interface navigation path, such as "device list-air conditioner control panel"; the fine-grained temporal feature dimension is 64, which characterizes the specific control operation, such as "click the temperature plus button".
[0079] Next, the importance of temporal features at different levels is weighted through a multi-head attention mechanism. First, based on high-level temporal features, three independent linear transformation layers are used to generate query matrices, key matrices, and value matrices, all with a matrix dimension of 256×256. The dot product of the query matrix and the key matrix is calculated to obtain the attention score, which is then normalized by softmax to obtain the high-level attention weight. The weight is multiplied by the value matrix to obtain the high-level context vector.
[0080] The high-level context vector is concatenated with the middle-level temporal features to form a 384-dimensional conditional query matrix, and a similar attention calculation process is repeated to obtain the middle-level context vector. The middle-level context vector is then concatenated with the fine-grained temporal features to form a 192-dimensional fine-grained conditional query matrix, and the fine-grained attention weights and context vectors are calculated.
[0081] Finally, hierarchical intent recognition is performed through a three-layer fully connected network. The first layer receives a 256-dimensional high-level context vector and outputs a probability distribution of 10 categories of high-level intents; the second layer outputs 20 categories of mid-level intent probabilities based on high-level intent probabilities and 128-dimensional mid-level context vectors; the third layer combines the results of the first two layers and a 64-dimensional fine-grained context vector to output a probability distribution of 50 categories of fine-grained operation intents.
[0082] In the training phase, the cross entropy loss between the prediction results and the true labels of the three levels is calculated respectively. At the same time, the hierarchical consistency constraint is introduced, requiring that the probability distribution of the high-level intent and the middle-level intent have a certain correlation, and the middle-level and fine-grained layers must also meet similar constraints. These loss terms are weighted and combined by weight coefficients of 0.4, 0.3, and 0.3 to construct the final optimization objective function. The Adam optimizer is used for model training, the learning rate is set to 0.001, the batch size is 64, and the training rounds are 100 rounds.
[0083] This application can achieve:
[0084] The bidirectional long short-term memory network is used to achieve bidirectional modeling of user operation behavior sequences, effectively capturing long-term and short-term temporal dependencies and improving the accuracy and robustness of intent recognition. The multi-head attention mechanism is used to achieve adaptive weighted fusion of temporal features at different levels, highlighting the contribution of important features, suppressing the interference of irrelevant features, and enhancing the feature expression ability of the model. Based on the hierarchical intent classification network, progressive intent prediction from coarse-grained to fine-grained is achieved. The introduction of hierarchical consistency constraints ensures the semantic consistency of intent prediction results at different levels, improving the interpretability of intent recognition.
[0085] In an optional implementation, similarity matching is performed in the prompt strategy library based on the user operation intention, and the prompt template with the highest matching degree is selected as the basic template; the operation probability distribution is used as the importance weight of the prompt content, and the prompt elements in the basic template are prioritized, including:
[0086] Constructing a multi-level prompt strategy library, wherein the multi-level prompt strategy library includes a plurality of strategy templates, each of which includes a template body text, a prompt element set, and an element dependency matrix, wherein the element dependency matrix is used to characterize the association relationship between each prompt element in the prompt element set;
[0087] Using a first semantic encoder to encode the user operation intention to obtain an intention semantic vector, using a second semantic encoder to encode the template body text of each of the policy templates in the multi-level prompt policy library to obtain a template semantic vector; calculating the cosine similarity between the intention semantic vector and the template semantic vector to obtain a first similarity component, calculating the correlation between the intention semantic vector and each prompt element in the prompt element set to obtain a second similarity component, performing a weighted combination of the first similarity component and the second similarity component to obtain a comprehensive similarity, selecting the best matching policy template from the multi-level prompt policy library based on the comprehensive similarity, and selecting the prompt template with the highest matching degree as the basic template;
[0088] Constructing an intention factor mapping matrix, performing matrix multiplication operation on the operation probability distribution and the intention factor mapping matrix to obtain an initial factor weight vector; performing weighted combination on the initial factor weight vector and the factor dependency matrix to obtain a fused factor weight vector, wherein the fused factor weight vector considers both intention relevance and factor dependency;
[0089] Based on the fusion element weight vector, the prompt element set in the optimal matching strategy template is sorted by importance to obtain an ordered prompt element sequence; according to the ordered prompt element sequence and the fusion element weight vector, the template body text of the optimal matching strategy template is reconstructed to generate importance-aware adaptive prompt content.
[0090] The prompt strategy optimization system based on user operation intention first builds a multi-level prompt strategy library. The strategy library uses a tree structure to organize multiple strategy templates. Each strategy template contains the template body text, prompt element set, and element dependency matrix. Taking the dialogue system as an example, the template body text can be "generate answers based on user questions", the prompt element set includes key elements such as question understanding, knowledge retrieval, and answer generation, and the element dependency matrix describes the correlation strength between these elements.
[0091] When performing template matching, the system uses a dual encoder architecture. The first semantic encoder uses a pre-trained language model to encode the user's operation intention and obtains a fixed-dimensional intent semantic vector. The second semantic encoder also uses a pre-trained language model, but is fine-tuned for the policy template to better capture template features. Taking the user query "how to improve code performance" as an example, the intent semantic vector will highlight the semantic features related to performance optimization.
[0092] The system calculates the similarity between the intent semantic vector and the template semantic vector, while considering the relevance of the intent to each prompt element. For example, the performance optimization intent has a high relevance to elements such as code analysis and performance measurement. By weightedly combining these two similarity components, the system can select the most suitable policy template.
[0093] Next, we construct an intent factor mapping matrix that describes the correspondence between the operation probability distribution and the prompt factors. The initial factor weights are obtained through matrix operations, and then weighted fusion is performed based on the factor dependencies. For example, the weight of the code analysis factor will be increased due to its strong dependency with the performance measurement factor.
[0094] Finally, the system sorts the prompt elements according to the weights of the fused elements and reconstructs the main text of the template accordingly. For performance optimization scenarios, the system will give priority to high-weight elements such as code analysis and performance bottleneck identification to generate more targeted prompt content.
[0095] For example, a multi-level prompt strategy library is first constructed, and a three-layer tree structure is used to organize strategy templates. Taking the intelligent customer service scenario as an example, the first layer is the task type, including consultation inquiry, fault handling, business handling, etc.; the second layer is the specific business scenario, such as product consultation, order inquiry, network failure, etc.; the third layer is the detailed prompt template. Each prompt template contains the template body text, prompt element set and element dependency matrix.
[0096] Taking the template of the "network fault handling" scenario as an example, the main text of the template is "Diagnose the cause of the network fault phenomenon described by the user and provide a solution." The prompt element set includes four core elements: fault phenomenon analysis, cause location, solution generation, and operation guidance. In the element dependency matrix, the correlation between fault phenomenon analysis and cause location is 0.8, the correlation between cause location and solution generation is 0.9, and the correlation between solution generation and operation guidance is 0.7.
[0097] After receiving the user's operation intention, the first semantic encoder based on BERT is used to encode the user's intention. Assuming that the user feedback is "the mobile WiFi cannot connect to the router", the encoder converts it into a 768-dimensional intention semantic vector. At the same time, the second semantic encoder is used to encode the main text of all templates in the policy library to obtain the template semantic vector of the same dimension.
[0098] The first similarity component is obtained by calculating the cosine similarity between the intention semantic vector and the template semantic vector. For the above user intention, the similarity with the "network fault handling" template is 0.85. At the same time, the correlation between the intention semantic vector and each prompt element is calculated to obtain the second similarity component. For example, the correlation of fault phenomenon analysis is 0.9, and the correlation of cause location is 0.8. The two similarity components are combined with weights of 0.6 and 0.4 to obtain a comprehensive similarity of 0.83.
[0099] Construct an intention factor mapping matrix that describes the correspondence between different operation intentions and prompt factors. Multiply the user's operation probability distribution by the mapping matrix to obtain the initial factor weight vector. In this example, the weight of fault phenomenon analysis is 0.9, the weight of cause location is 0.8, the weight of solution generation is 0.7, and the weight of operation guidance is 0.6.
[0100] The initial factor weight vector and the factor dependency matrix are weighted and fused, with weights of 0.7 and 0.3 respectively, to obtain the fused factor weight vector. After fusion, the weight of fault phenomenon analysis is 0.85, the weight of cause location is 0.82, the weight of solution generation is 0.75, and the weight of operation guidance is 0.65.
[0101] Based on the fusion factor weight vector, the prompt elements are ranked in importance to obtain an ordered sequence of prompt elements: fault phenomenon analysis, cause location, solution generation, and operation guidance. The template body text is reconstructed according to the sequence and weight vector to generate importance-aware adaptive prompt content: "Please first analyze the WiFi connection fault phenomenon described by the user in detail (pay attention to signal strength, authentication status, etc.), then locate the possible fault cause based on the phenomenon, generate a targeted solution, and finally provide clear operation guidance steps."
[0102] This application can achieve:
[0103] Improved the accuracy and adaptability of the prompt strategy. Through multi-dimensional similarity matching and element importance ranking, the system can more accurately understand the user's intention and generate corresponding prompt content. Enhanced the structured degree of the prompt content. The weight fusion mechanism based on element dependencies makes the generated prompts more logically coherent and the relationship between the elements closer. Dynamic optimization of the prompt strategy is achieved. The system can adaptively adjust the importance of the prompt elements according to the changes in the user's operation intention, improving the flexibility and efficiency of the system response.
[0104] In an optional implementation, a reinforcement learning method is used to continuously optimize the prompt generation strategy, and the interactive feedback data of the user on the historical prompt information is used as a reward signal. The dynamic adjustment rules of the prompt content are learned through the strategy network; the basic template is rewritten and supplemented according to the dynamic adjustment rules to generate personalized prompt information that conforms to the current context, including:
[0105] Construct a multi-dimensional reward function for user feedback, wherein the multi-dimensional reward function includes a click response reward, a stay time reward, a task completion reward, and a satisfaction score reward, and calculate the user's interactive feedback score for the historical prompt information based on the multi-dimensional reward function;
[0106] Acquire current context features, historical interaction sequence codes, prompt template features, and user portrait features, and combine the current context features, the historical interaction sequence codes, the prompt template features, and the user portrait features to generate a state vector; construct an action space, the action space including modification operations, append operations, deletion operations, and reordering operations, and each operation includes an operation type parameter, a location information parameter, and a content feature parameter;
[0107] Constructing a policy network and a value network, wherein the policy network is used to predict the execution probability distribution of each operation according to the state vector, and the value network is used to evaluate the value score of the state vector; calculating the advantage function value based on the interactive feedback score and the value score, performing policy gradient update on the parameters of the policy network according to the advantage function value, and performing value estimation update on the parameters of the value network according to the temporal difference error;
[0108] Constructing a prompt adjustment rule set based on the updated policy network, wherein each rule in the prompt adjustment rule set includes a trigger condition, an adjustment action, and a rule priority; calculating an execution probability of each adjustment action according to the state vector, and multiplying the execution probability by a user personalized weight to obtain a comprehensive execution priority;
[0109] Based on the comprehensive execution priority, an optimal adjustment action is selected from the action space, and the prompt template is dynamically adjusted according to the optimal adjustment action to generate personalized prompt content; the user's interactive feedback data on the personalized prompt content is input into the multi-dimensional reward function, and the parameters of the policy network and the value network are continuously optimized to realize online learning of the prompt strategy.
[0110] First, a multi-dimensional reward function system is constructed. By collecting the user's click behavior data on the prompt information, recording the user's stay time on the prompt content, collecting the task completion mark data, and the user's score data on the prompt effect. The basic score for the click response is set to 10 points, and each valid click is accumulated by 2 points; the stay time is accumulated at 1 point every 10 seconds; the task completion is given an evaluation score of 0-100 points based on the achievement of the preset task nodes; the satisfaction score adopts a five-star system, with each star corresponding to 20 points. The score data of these four dimensions are normalized and weighted summed to obtain the comprehensive interactive feedback score.
[0111] Then, the feature vector is constructed and combined. Semantic features, intention features, and emotional features are extracted from the current conversation to form a context feature vector; historical interaction records are converted into a fixed-dimensional sequence feature vector through a sequence encoder; structural features and theme features of the template are obtained from the prompt template library to form a template feature vector; user basic attributes, behavioral habits, preferences, and other information are collected to construct a user portrait feature vector. These feature vectors are spliced and reduced in dimension to generate the final state vector.
[0112] In the design of the action space, the modification operation includes parameters such as replacing words, adjusting the tone, changing the sentence structure, etc.; the append operation includes parameters such as supplementary explanation, example explanation, and related recommendation; the deletion operation includes parameters such as redundant information filtering and duplicate content removal; the reordering operation includes parameters such as importance sorting, time sequence sorting, and logical sorting. Each operation is configured with corresponding location index information and content feature description.
[0113] The policy network uses a multi-layer neural network structure. The input layer receives the state vector, the middle layer extracts and combines features, and the output layer generates the probability distribution of each action. The value network also uses a similar network structure, but the output layer only generates a value evaluation score. Based on the user interaction feedback score and the value network evaluation score, the policy gradient is calculated and the network parameters are updated.
[0114] According to the trained policy network, the weight parameters and decision rules in the network are extracted to build a prompt adjustment rule library. Each rule contains a specific trigger condition description, specific adjustment action instructions and initial priority settings. In actual application, the current state vector is input into the policy network to obtain the execution probability of each action, and then combined with the user's personalized weight parameters to calculate the final execution priority.
[0115] Finally, the adjustment action with the highest priority is selected to modify, append, delete or reorder the basic prompt template accordingly to generate personalized prompt content suitable for the current user and scenario. The system continuously collects new feedback data from users, continuously optimizes the parameters of the policy network and value network, and realizes dynamic adjustment of the prompt generation strategy.
[0116] For example, we first construct a multi-dimensional reward function to evaluate user feedback. Taking the intelligent customer service scenario as an example, the click response reward is calculated based on the user's behavior of clicking on the recommended link in the prompt content, and one click is counted as 1 point; the stay time reward is calculated based on the time the user reads the prompt content, and every 10 seconds of stay is counted as 0.5 points; the task completion reward is calculated based on whether the user successfully solves the problem according to the prompt, and a complete solution is counted as 5 points, and a partial solution is counted as 3 points; the satisfaction score reward is based on the user's five-star evaluation conversion, 5 points for five stars, 4 points for four stars, and so on.
[0117] In actual applications, if a user receives a network troubleshooting prompt, clicks on the recommended troubleshooting link (1 point), stays to read for 40 seconds (2 points), partially solves the problem after following the prompt (3 points), and finally gives a four-star rating (4 points), the total reward score for this interaction is 10 points. By accumulating the reward score for each prompt in the historical interaction sequence, a prompt effect evaluation benchmark is established.
[0118] Obtain state representation information, including: current context features (such as the user's most recent 3 operations, current interface status, etc.), historical interaction sequence encoding (converting the most recent 10 interaction records into feature vectors through a sequence encoder), prompt template features (template type, key elements, semantic vectors, etc.), user portrait features (usage habits, preference settings, historical data, etc.). These features are concatenated to form a 512-dimensional state vector.
[0119] Construct an action space to define optional adjustment operations: modify operation (replace specific content in the template, such as changing "check network settings" to "check WiFi settings"), append operation (insert additional instructions at a specific location), delete operation (remove redundant or irrelevant content), reorder operation (adjust the display order of content modules). Each operation contains parameter information such as operation type, location index, and specific content.
[0120] Establish the policy network and value network structure. The policy network uses a three-layer fully connected network. The 512 nodes in the input layer receive the state vector, the 256 nodes in the hidden layer use the ReLU activation function, and the output layer uses the Softmax function to obtain the operation probability distribution corresponding to the action space dimension. The value network structure is similar, but the output layer has only one node, which is used to predict the state value score.
[0121] Taking the processing of network failure prompts as an example, after receiving the current state input, the policy network may predict the following operation probabilities: the probability of changing "network settings" to "WiFi settings" is 0.4, the probability of adding "Please confirm the status of the router indicator light" is 0.3, the probability of deleting "Check the network cable connection" (irrelevant) is 0.2, and the probability of adjusting the sorting is 0.1. The value network also predicts that the value score of the current state is 7.5.
[0122] The advantage function value is calculated based on the difference between the actual interaction feedback score and the value network prediction score. If the actual score is 10 points and the predicted score is 7.5 points, the advantage value is 2.5 points. This advantage value is used to guide the policy network parameter update and increase the probability of generating effective adjustment operations. At the same time, the time difference error is used to optimize the value network to improve the accuracy of state value prediction.
[0123] Based on the updated policy network, action selection rules are extracted. For example, when it is detected that the user is using mobile WiFi, the rule of replacing "network settings" with "WiFi settings" is triggered, and the rule priority is 0.8; when the user views the same prompt multiple times, the rule of adding additional instructions is triggered, and the priority is 0.6. The personalized weight is calculated based on the user's historical operation habits. If the user prefers concise prompts, the weight of the delete operation is 1.2, and the weight of the add operation is 0.8.
[0124] Multiply the operation probability by the personalized weight to get the final execution priority. Select the highest priority adjustment operation to perform dynamic optimization of the prompt content. Continue to collect user feedback on the optimized prompt content, input the reward function for evaluation, continuously iterate and optimize the policy network and value network parameters, and realize adaptive learning of the prompt strategy.
[0125] This application can achieve:
[0126] Through the design and application of multi-dimensional reward functions, the user's acceptance of the prompt content and the effect of its use are comprehensively evaluated, making the optimization direction of the prompt strategy more accurate, and significantly improving the quality and pertinence of the prompt content. The deep reinforcement learning method is used to achieve continuous optimization of the prompt strategy. The system can automatically learn rules and patterns from user interaction data, and the intelligent level of prompt generation is continuously improved, reducing the workload of manual adjustment. The personalized prompt generation mechanism based on user portraits and contextual features enables the prompt content to better meet the personalized needs of different users, improve user experience and satisfaction, and at the same time significantly improve the timeliness and scene adaptability of the prompt content.
[0127] In an optional implementation, a prompt adjustment rule set is constructed based on the updated policy network, each rule in the prompt adjustment rule set includes a trigger condition, an adjustment action and a rule priority; the execution probability of each adjustment action is calculated according to the state vector, and the execution probability is multiplied by the user personalized weight to obtain a comprehensive execution priority including:
[0128] Constructing a rule representation model, wherein the rule representation model includes a rule trigger condition set, an adjustment action description, a rule priority, and a rule credibility, wherein each trigger condition in the rule trigger condition set includes a feature attribute, a comparison operator, and a threshold parameter, and the rule trigger condition set is used to filter the state vector;
[0129] The state vector is processed by using a policy network to generate candidate actions and their corresponding action probabilities, the gradient information of the state vector relative to each candidate action is calculated, and the state-action pair whose action probability is greater than a preset probability threshold and whose gradient significance is greater than a preset gradient threshold is used as an initial adjustment rule;
[0130] The initial adjustment rule is evaluated, rule confidence is calculated based on historical samples that meet the rule triggering condition set, rule support is calculated based on the historical triggering times of the initial adjustment rule, rule novelty is calculated based on the difference between the initial adjustment rule and the existing rule set, and rule confidence, rule support and rule novelty are weightedly combined to obtain rule priority;
[0131] Acquire user interaction data, extract multi-dimensional user features from the user interaction data to construct a user preference vector, map the candidate actions to a feature space using an action mapping function, map the user preference vector to the same feature space, and calculate a personalized weight based on the matching degree of the two mapping results;
[0132] Multiplying the action probability by the personalized weight and introducing a decay factor based on the complexity of the action to obtain a comprehensive execution priority of each of the candidate actions, wherein the complexity of the action is determined according to the number of operation steps involved in the action;
[0133] The rule set is dynamically maintained based on the comprehensive execution priority and the rule priority, a final priority score of each rule in the rule set is calculated, rules whose final priority scores are higher than a priority threshold and whose usage frequencies are higher than a usage frequency threshold are retained in the rule set, and a personalized prompt adjustment strategy is generated based on the rule set.
[0134] First, a rule representation model is constructed, which adopts a multi-layer structure design. The rule trigger condition set contains multiple feature dimensions, such as user behavior features, environmental features, and interaction scene features. Each trigger condition is defined by feature attributes, comparison operators, and threshold parameters, such as "user clicks are greater than 10", "session duration is more than 5 minutes", etc. For the input state vector, these trigger conditions are used to filter and determine which rules may be activated.
[0135] Then, the state vector is processed using a deep policy network to generate candidate actions and their probability distribution. Taking the user interaction scenario as an example, the state vector contains information such as the user's current behavior characteristics and historical interaction records, and the network outputs possible prompt adjustment actions, such as "simplify prompt language" and "add example description". By calculating the sensitivity of the state vector to each candidate action, the state-action pair with an action probability greater than 0.6 and a gradient significance greater than 0.3 is selected as the initial adjustment rule.
[0136] Perform a multi-dimensional evaluation of the initial adjustment rules. Calculate the rule confidence based on historical data, such as the proportion of positive effects produced after the rule was triggered in the past; calculate the rule support, that is, the proportion of the frequency of the rule being triggered in the total sample; calculate the rule novelty, measured by the semantic similarity with the existing rule set. Weight these three indicators in a ratio of 4:3:3 to obtain the rule priority.
[0137] User features are extracted from user interaction data, including preferred topics, expressions, interaction frequency and other dimensions, to construct a user preference vector. Candidate actions are converted into vector representations of the same dimension through feature mapping, and the cosine similarity with the user preference vector is calculated as the personalized weight. For example, if the user prefers concise expression, the action of simplifying the prompt will receive a higher weight.
[0138] The calculation of the comprehensive execution priority takes into account the complexity of the action, which is determined by the steps involved in the action. For example, the complexity of "adjusting the tone of voice" is 1, and the complexity of "reconstructing the prompt structure" is 3. The final priority is the product of the action probability, the personalization weight, and the complexity attenuation factor.
[0139] Dynamically maintain the rule set based on comprehensive execution priority and rule priority. Calculate the final priority score of the rules, and retain the rules with a score higher than 0.7 and an average monthly usage frequency of more than 100 times. The resulting personalized prompt adjustment strategy can adapt to the needs of different users.
[0140] For example, a rule representation model is constructed and the basic structure of the rule is designed. Taking the intelligent customer service scenario as an example, the rule trigger condition set includes user behavior characteristics, interface status characteristics, and time characteristics. For the "network fault handling" scenario, the trigger conditions may include: the number of times the user submits similar questions continuously is greater than 3 times, the problem description contains the keyword "WiFi connection failure", and the page stay time in the last 10 minutes is less than 30 seconds. Each trigger condition is defined by feature attributes (such as the number of submissions), comparison operators (such as greater than), and threshold parameters (such as 3 times).
[0141] The description of the adjustment action includes the action type, the operation object and the specific parameters. For example, the action type can be "replace text", the operation object is a specific paragraph in the prompt content, and the parameter is the new content after replacement. The rule priority uses a value between 0 and 1 to indicate the importance of the rule, and the rule credibility reflects the reliability of the rule.
[0142] Use the trained policy network to process the current state vector. Take the user feedback "the phone cannot connect to WiFi" as an example, convert the user description, historical operation records, interface status and other information into a 512-dimensional state vector and input it into the policy network. The network outputs multiple candidate actions and their probability distribution: for example, the action probability of "add WiFi troubleshooting steps" is 0.8, the action probability of "simplify professional terms" is 0.6, and the action probability of "add graphic description" is 0.5.
[0143] At the same time, the gradient information of the state vector for each candidate action is calculated to indicate the sensitivity of the action to the current state. The preset probability threshold is set to 0.6 and the gradient threshold is set to 0.3. The two actions of "adding WiFi troubleshooting steps" and "simplifying professional terms" are selected as the initial adjustment rules.
[0144] Perform a multi-dimensional evaluation of the initial adjustment rules. Assuming that the "Add WiFi troubleshooting steps" rule has a user problem solving rate of 85% after being triggered in historical samples, the rule confidence is calculated to be 0.85; the number of historical rule triggers accounts for 30% of the total number of rule triggers, and the rule support is 0.3; the average text similarity with the rules in the existing rule library is 0.4, and the rule novelty is 0.6.
[0145] The confidence, support and novelty of the rule are weighted according to the weights of 0.5, 0.3 and 0.2, and the priority of the rule is 0.66. Similarly, the priority of the "Simplify Professional Terms" rule is 0.58.
[0146] Analyze user historical interaction data, extract user features and construct preference vectors. Feature dimensions include: preference for graphic presentation (score 0.8), preference for concise description (score 0.9), professional knowledge level (score 0.3), etc. Use the action mapping function to convert candidate actions to the same feature space, and calculate the matching degree with user preferences as the personalized weight.
[0147] The personalized weight of the "Add WiFi troubleshooting steps" action is 0.7. Considering that this action contains 3 operation steps, the complexity attenuation factor of 0.9 is introduced. The final action probability of 0.8 is multiplied by the personalized weight of 0.7 and the attenuation factor is considered to obtain a comprehensive execution priority of 0.504. Similarly, the comprehensive execution priority of the "Simplify professional terms" action is calculated to be 0.486.
[0148] The priority threshold is set to 0.5 and the usage frequency threshold is set to 0.2. The rule set is filtered based on the comprehensive execution priority and rule priority. The final priority score of the "Add WiFi troubleshooting steps" rule exceeds the threshold and the usage frequency meets the requirements, so it is retained in the rule set. Finally, a personalized prompt policy is generated based on the retained rules: "Please follow the steps below to check the WiFi connection: 1. Confirm that the WiFi switch on the mobile phone is turned on; 2. Check the list of available networks; 3. Enter the correct network password."
[0149] This application can achieve:
[0150] Through multi-dimensional evaluation and dynamic maintenance mechanisms, the accuracy and timeliness of prompt adjustment rules are improved, enabling the system to continuously optimize and maintain the high quality of the rule base. The introduction of user personalized weights and action complexity attenuation factors enables personalized customization of prompt adjustment strategies, improving user experience while ensuring system efficiency. The use of a deep policy network-based method to automatically generate adjustment rules reduces the cost of manual rule design while improving the efficiency and quality of rule generation.
[0151] Figure 2FIG. 1 is a schematic diagram of a system for automatically optimizing software operation prompts based on context understanding according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0152] The first unit is used to collect and construct a multi-dimensional context feature data set of user software operations, wherein the multi-dimensional context feature data set includes at least one of historical operation sequence data of the user within a preset time window, interface state data of the current software, real-time user interaction behavior data, and software operating environment data; and perform data standardization and time sequence alignment processing on the multi-dimensional context feature data set to generate pre-processed context feature data;
[0153] The second unit is used to input the preprocessed context feature data into a multimodal feature fusion network, perform feature encoding on historical operation sequence data, interface state data, user real-time interaction behavior data and software operating environment data, and generate a multimodal feature vector; use a bidirectional long short-term memory network to perform temporal dependency modeling on the multimodal feature vector, capture the dynamic evolution law of user operation behavior, and output temporal context features; based on a multi-head attention mechanism, the importance of different dimensions of the temporal context features is weighted to generate a fused context semantic representation vector; the context semantic representation vector is input into an intention classification network to predict user operation intentions and operation probability distribution;
[0154] The third unit is used to perform similarity matching in the prompt strategy library based on the user operation intention, and select the prompt template with the highest matching degree as the basic template; use the operation probability distribution as the importance weight of the prompt content, and prioritize the prompt elements in the basic template; use the reinforcement learning method to continuously optimize the prompt generation strategy, use the user's interactive feedback data on historical prompt information as a reward signal, and learn the dynamic adjustment rules of the prompt content through the strategy network; rewrite and supplement the basic template according to the dynamic adjustment rules to generate personalized prompt information that conforms to the current context.
[0155] According to a third aspect of the embodiments of the present invention,
[0156] An electronic device is provided, comprising:
[0157] processor;
[0158] a memory for storing processor-executable instructions;
[0159] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0160] According to a fourth aspect of the embodiments of the present invention,
[0161] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0162] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatically optimizing software operation prompts based on context understanding, characterized in that: include: Collect and construct a multi-dimensional context feature data set of user software operations, wherein the multi-dimensional context feature data set includes at least one of historical operation sequence data of the user within a preset time window, interface state data of the current software, real-time user interaction behavior data, and software operating environment data; perform data standardization and time sequence alignment processing on the multi-dimensional context feature data set to generate pre-processed context feature data; Inputting the preprocessed context feature data into a multimodal feature fusion network, performing feature encoding on the historical operation sequence data, the interface state data, the user's real-time interactive behavior data, and the software running environment data, respectively, to generate a multimodal feature vector; A bidirectional long short-term memory network is used to model the temporal dependency of the multimodal feature vector, capture the dynamic evolution law of the user's operation behavior, and output the temporal context features; based on the multi-head attention mechanism, the importance of different dimensions of the temporal context features is weighted to generate a fused context semantic representation vector; the context semantic representation vector is input into the intention classification network to predict the user's operation intention and operation probability distribution; Based on the user operation intention, similarity matching is performed in the prompt strategy library, and the prompt template with the highest matching degree is selected as the basic template; Using the operation probability distribution as the importance weight of the prompt content, and prioritizing the prompt elements in the basic template; The reinforcement learning method is used to continuously optimize the prompt generation strategy. The user's interactive feedback data on historical prompt information is used as a reward signal. The dynamic adjustment rules of the prompt content are learned through the strategy network. The basic template is rewritten and supplemented according to the dynamic adjustment rules to generate personalized prompt information that conforms to the current context.
2. The method according to claim 1, characterized in that The historical operation sequence data records the access order of function modules and the timestamp of operation events; the interface status data of the current software includes the hierarchical structure information of the interface function modules and the spatial distribution information of the interactive controls; the real-time user interactive behavior data includes the coordinate sequence of the mouse movement trajectory, the operation dwell time statistics and the control triggering frequency; the software operating environment data includes the software version identification information, system configuration parameters and hardware equipment information.
3. The method according to claim 1, characterized in that A bidirectional long short-term memory network is used to model the temporal dependency of the multimodal feature vector, capture the dynamic evolution law of the user's operation behavior, and output the temporal context feature; based on the multi-head attention mechanism, the importance of different dimensions of the temporal context feature is weighted to generate a fused context semantic representation vector; Inputting the context semantic representation vector into the intention classification network, predicting the user operation intention and operation probability distribution includes: Inputting the multimodal feature vector into a bidirectional long short-term memory network, the bidirectional long short-term memory network includes a forward propagation layer and a backward propagation layer, and performing time series modeling on the multimodal feature vector through the forward propagation layer and the backward propagation layer, respectively, to obtain high-level time series features, middle-level time series features and fine-grained time series features, wherein the high-level time series features are used to characterize function module selection information, the middle-level time series features are used to characterize interface navigation information, and the fine-grained time series features are used to characterize control operation information; Generate a query matrix, a key matrix and a value matrix based on the high-level temporal features, calculate the similarity between the query matrix and the key matrix to obtain a high-level attention weight, and weight the high-level attention weight and the value matrix to obtain a high-level context vector; splice the high-level context vector with the middle-level temporal features to generate a conditional query matrix, calculate the middle-level attention weight based on the conditional query matrix and obtain the middle-level context vector; splice the middle-level context vector with the fine-grained temporal features to generate a fine-grained conditional query matrix, calculate the fine-grained attention weight based on the fine-grained conditional query matrix and obtain the fine-grained context vector; Input the high-level context vector into the first fully connected layer of the intent classification network to obtain a high-level intent probability distribution; based on the combined features of the high-level intent probability distribution and the middle-level context vector, obtain the middle-level intent probability distribution through the second fully connected layer; based on the combined features of the high-level intent probability distribution, the middle-level intent probability distribution and the fine-grained context vector, obtain the fine-grained intent probability distribution through the third fully connected layer; Calculate the cross entropy loss of the high-level intent probability distribution, the mid-level intent probability distribution and the fine-grained intent probability distribution relative to the true label, and calculate the relative entropy between adjacent level intent probability distributions to obtain the hierarchical consistency constraint term; perform a weighted combination of the cross entropy loss and the hierarchical consistency constraint term to construct a joint loss function, and optimize the parameters of the bidirectional long short-term memory network, the first fully connected layer, the second fully connected layer and the third fully connected layer through the gradient descent algorithm.
4. The method according to claim 1, characterized in that: Based on the user operation intention, similarity matching is performed in the prompt strategy library, and the prompt template with the highest matching degree is selected as the basic template; Using the operation probability distribution as the importance weight of the prompt content, and prioritizing the prompt elements in the basic template includes: Constructing a multi-level prompt strategy library, wherein the multi-level prompt strategy library includes a plurality of strategy templates, each of which includes a template body text, a prompt element set, and an element dependency matrix, wherein the element dependency matrix is used to characterize the association relationship between each prompt element in the prompt element set; Using a first semantic encoder to encode the user operation intention to obtain an intention semantic vector, using a second semantic encoder to encode the template body text of each of the policy templates in the multi-level prompt policy library to obtain a template semantic vector; calculating the cosine similarity between the intention semantic vector and the template semantic vector to obtain a first similarity component, calculating the correlation between the intention semantic vector and each prompt element in the prompt element set to obtain a second similarity component, performing a weighted combination of the first similarity component and the second similarity component to obtain a comprehensive similarity, selecting the best matching policy template from the multi-level prompt policy library based on the comprehensive similarity, and selecting the prompt template with the highest matching degree as the basic template; Constructing an intention factor mapping matrix, performing matrix multiplication operation on the operation probability distribution and the intention factor mapping matrix to obtain an initial factor weight vector; performing weighted combination on the initial factor weight vector and the factor dependency matrix to obtain a fused factor weight vector, wherein the fused factor weight vector considers both intention relevance and factor dependency; Based on the fusion element weight vector, the prompt element set in the optimal matching strategy template is sorted by importance to obtain an ordered prompt element sequence; according to the ordered prompt element sequence and the fusion element weight vector, the template body text of the optimal matching strategy template is reconstructed to generate importance-aware adaptive prompt content.
5. The method according to claim 1, characterized in that The reinforcement learning method is used to continuously optimize the prompt generation strategy, and the user's interactive feedback data on historical prompt information is used as a reward signal. The dynamic adjustment rules of prompt content are learned through the strategy network; Rewriting and supplementing the basic template according to the dynamic adjustment rule to generate personalized prompt information that conforms to the current context includes: Construct a multi-dimensional reward function for user feedback, wherein the multi-dimensional reward function includes a click response reward, a stay time reward, a task completion reward, and a satisfaction score reward, and calculate the user's interactive feedback score for the historical prompt information based on the multi-dimensional reward function; Acquire current context features, historical interaction sequence codes, prompt template features, and user portrait features, and combine the current context features, the historical interaction sequence codes, the prompt template features, and the user portrait features to generate a state vector; construct an action space, the action space including modification operations, append operations, deletion operations, and reordering operations, and each operation includes an operation type parameter, a location information parameter, and a content feature parameter; Constructing a policy network and a value network, wherein the policy network is used to predict the execution probability distribution of each operation according to the state vector, and the value network is used to evaluate the value score of the state vector; calculating the advantage function value based on the interactive feedback score and the value score, performing policy gradient update on the parameters of the policy network according to the advantage function value, and performing value estimation update on the parameters of the value network according to the temporal difference error; Constructing a prompt adjustment rule set based on the updated policy network, wherein each rule in the prompt adjustment rule set includes a trigger condition, an adjustment action, and a rule priority; calculating an execution probability of each adjustment action according to the state vector, and multiplying the execution probability by a user personalized weight to obtain a comprehensive execution priority; Based on the comprehensive execution priority, an optimal adjustment action is selected from the action space, and the prompt template is dynamically adjusted according to the optimal adjustment action to generate personalized prompt content; the user's interactive feedback data on the personalized prompt content is input into the multi-dimensional reward function, and the parameters of the policy network and the value network are continuously optimized to realize online learning of the prompt strategy.
6. The method according to claim 5, characterized in that Building a prompt adjustment rule set based on the updated policy network, each rule in the prompt adjustment rule set includes a trigger condition, an adjustment action, and a rule priority; Calculating the execution probability of each adjustment action according to the state vector, and multiplying the execution probability by the user personalized weight to obtain a comprehensive execution priority includes: Constructing a rule representation model, wherein the rule representation model includes a rule trigger condition set, an adjustment action description, a rule priority, and a rule credibility, wherein each trigger condition in the rule trigger condition set includes a feature attribute, a comparison operator, and a threshold parameter, and the rule trigger condition set is used to filter the state vector; The state vector is processed by using a policy network to generate candidate actions and their corresponding action probabilities, the gradient information of the state vector relative to each candidate action is calculated, and the state-action pair whose action probability is greater than a preset probability threshold and whose gradient significance is greater than a preset gradient threshold is used as an initial adjustment rule; The initial adjustment rule is evaluated, rule confidence is calculated based on historical samples that meet the rule triggering condition set, rule support is calculated based on the historical triggering times of the initial adjustment rule, rule novelty is calculated based on the difference between the initial adjustment rule and the existing rule set, and rule confidence, rule support and rule novelty are weightedly combined to obtain rule priority; Acquire user interaction data, extract multi-dimensional user features from the user interaction data to construct a user preference vector, map the candidate actions to a feature space using an action mapping function, map the user preference vector to the same feature space, and calculate a personalized weight based on the matching degree of the two mapping results; Multiplying the action probability by the personalized weight and introducing a decay factor based on the complexity of the action to obtain a comprehensive execution priority of each of the candidate actions, wherein the complexity of the action is determined according to the number of operation steps involved in the action; The rule set is dynamically maintained based on the comprehensive execution priority and the rule priority, a final priority score of each rule in the rule set is calculated, rules whose final priority scores are higher than a priority threshold and whose usage frequencies are higher than a usage frequency threshold are retained in the rule set, and a personalized prompt adjustment strategy is generated based on the rule set.
7. A software operation prompt automatic optimization system based on context understanding, used to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is used to collect and construct a multi-dimensional context feature data set of user software operations, wherein the multi-dimensional context feature data set includes at least one of historical operation sequence data of the user within a preset time window, interface state data of the current software, real-time user interaction behavior data, and software operating environment data; and perform data standardization and time sequence alignment processing on the multi-dimensional context feature data set to generate pre-processed context feature data; The second unit is used to input the preprocessed context feature data into a multimodal feature fusion network, perform feature encoding on the historical operation sequence data, the interface state data, the user's real-time interactive behavior data and the software running environment data, and generate a multimodal feature vector; A bidirectional long short-term memory network is used to model the temporal dependency of the multimodal feature vector, capture the dynamic evolution law of the user's operation behavior, and output the temporal context features; based on the multi-head attention mechanism, the importance of different dimensions of the temporal context features is weighted to generate a fused context semantic representation vector; the context semantic representation vector is input into the intention classification network to predict the user's operation intention and operation probability distribution; A third unit is used to perform similarity matching in a prompt strategy library based on the user operation intention, and select a prompt template with the highest matching degree as a basic template; Using the operation probability distribution as the importance weight of the prompt content, and prioritizing the prompt elements in the basic template; The reinforcement learning method is used to continuously optimize the prompt generation strategy. The user's interactive feedback data on historical prompt information is used as a reward signal. The dynamic adjustment rules of the prompt content are learned through the strategy network. The basic template is rewritten and supplemented according to the dynamic adjustment rules to generate personalized prompt information that conforms to the current context.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Long text user opinion understanding method and system based on memory enhancement
CN119293148A
Live broadcast room content identification and intelligent distribution method and system based on multi-modal fusion
CN119377895A
Cited By
Apparatus and method for generating context-aware device prompts and transmission protocols
US12683918B1