Operation guiding method and device, equipment and medium
By fusing user terminal question-and-answer, behavior, and permission information to generate feature vectors, and using a deep Q-network model to filter operation guidance, the problem of insufficient personalization in traditional methods is solved, and dynamic adaptation of personalized operation guidance is achieved, thereby improving operation efficiency and customer satisfaction.
Patent Information
- Application Number
- CN202510889887.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-04
AI Technical Summary
Existing operating instructions lack personalization, making it difficult to meet users' individual needs, resulting in low operating efficiency and low customer satisfaction.
By fusing question-and-answer information, behavioral information, and permission information from user terminals, a fused feature vector is generated. Then, a prediction model trained with a deep Q-network is used to dynamically filter and display personalized operation guidance.
It enables dynamic adaptation to different users' usage habits and permission levels based on real-time user interaction data, providing users with personalized operation guidance that meets their current needs, thereby improving operational efficiency and customer satisfaction.
Smart Images

Figure CN120892121A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and specifically to an operation guidance method, apparatus, device, and medium. Background Technology
[0002] With the acceleration of globalization and digital transformation, multinational corporations' LTC (Lead to Cash) systems are becoming increasingly complex, encompassing multiple aspects such as sales, marketing, customer service, finance, and risk. This complexity places higher demands on operators, but operators often have limited familiarity with the new systems, leading to low operational efficiency and impacting business flow and customer satisfaction.
[0003] Traditional methods typically generate operation instructions based on preset templates, suitable for standardized operation scenarios. However, different users may have significantly different learning habits and needs, and traditional methods lack personalization, making it difficult to meet users' individual requirements. Summary of the Invention
[0004] The purpose of this application is to provide an operation guidance method, apparatus, device, and medium to solve the problem that existing operation guidance methods lack personalization and are difficult to meet users' personalized needs.
[0005] To achieve the above objectives, the first aspect of this application provides an operational guidance method, the method comprising: The first question-and-answer information, first behavior information, and first permission information corresponding to the user terminal are fused to obtain a fused feature vector. The first question-and-answer information includes the question text and answer text generated by the user through the user terminal; the first behavior information includes the page click data of the user on the user terminal; and the first permission information includes the permission level of the user on the user terminal. The fused feature vectors are input into the pre-acquired prediction model, and the probability of multiple functional items is determined by the prediction model. Based on the order of probability from high to low among multiple functional items, at least one target functional item is selected from the multiple functional items. Display at least one target function on the user terminal's display interface.
[0006] In one embodiment of this application, the first question-and-answer information, the first behavior information, and the first permission information corresponding to the user terminal are fused to obtain a fused feature vector, including: Keyword extraction is performed on both the question text and the answer text, and the extracted keywords are converted into text feature vectors using a pre-trained word vector model; The page click data is arranged according to time series and processed by a recurrent neural network to obtain behavioral feature vectors; The permission levels are numerically encoded to obtain permission feature vectors; The text feature vector, behavior feature vector, and permission feature vector are concatenated to obtain the fused feature vector.
[0007] In one embodiment of this application, after displaying at least one target function item on the display interface of a user terminal, the method further includes: When a user clicks on any target function item, the system records the second question and answer information, the second behavior information, and the second permission information corresponding to that click. At preset time intervals, the prediction model is updated using the second question-and-answer information, the second behavior information, and the second permission information.
[0008] In one embodiment of this application, the prediction model is a model trained based on a deep Q-network, and the training process includes: Construct an experience replay buffer, which includes multiple experience data, each of which includes the user's corresponding current state, action, reward value, and next state. The first-depth Q-network was trained using multiple empirical datasets to obtain the predicted Q-values; The second deep Q-network was trained using multiple empirical data sets to obtain the target Q-value. The second deep Q-network and the first deep Q-network have the same network structure but different network parameters. The loss function value is determined based on the predicted Q value and the target Q value; If the loss function value does not meet the preset training stopping condition, return to the previous step and train the first deep Q network using multiple empirical data to obtain the predicted Q value until the loss function value meets the preset training stopping condition, at which point the first deep Q network is determined as the prediction model.
[0009] In one embodiment of this application, after determining the loss function based on the predicted Q-value and the target Q-value, the method further includes: At preset intervals, the model parameters of the first deep Q-network are copied to the second deep Q-network.
[0010] In one embodiment of this application, the reward value is determined through the following process: Based on preset evaluation indicators and recommended function items corresponding to actions, the interaction value of users to actions is calculated. The recommended function items are the function items predicted by the prediction model based on the current state. The reward value is determined based on the interaction value; The evaluation metrics include at least one of the following: click behavior reward, feature completion reward, dwell time reward, and penalty.
[0011] In one embodiment of this application, the loss function value is calculated using the following formula:
[0012] Where L is the loss function value, N It is the amount of empirical data. It is the i-th predicted Q value. It is the i-th objective Q value.
[0013] A second aspect of this application provides an operation guidance device, the device comprising: The fusion module is used to fuse the first question-and-answer information, the first behavior information, and the first permission information corresponding to the user terminal to obtain a fused feature vector. The first question-and-answer information includes the question text and answer text generated by the user through the user terminal; the first behavior information includes the page click data of the user on the user terminal; and the first permission information includes the user's permission level on the user terminal. The output module is used to input the fused feature vector into the pre-acquired prediction model, and the prediction model determines the probability of multiple functional items. The filtering module is used to filter out at least one target function item from multiple function items based on the order of their probabilities from high to low. The display module is used to display at least one target function item on the display interface of the user terminal.
[0014] A third aspect of this application provides an electronic device, which includes: a processor and a memory storing computer program instructions; When the processor executes computer program instructions, it implements the operation guidance method as described in any of the first aspects.
[0015] A third aspect of this application provides a machine-readable storage medium storing instructions that cause a machine to perform an operation instruction method according to any one of the first aspects.
[0016] The operation guidance method provided in this application generates a fused feature vector by fusing the first question-and-answer information, first action information, and first permission information of the user terminal. This comprehensively captures the user's personalized needs and operation scenarios. The fused feature vector is input into a prediction model to output the probability of function items and then filtered and displayed, changing the traditional method of generating operation guidance based on preset templates. Compared to the lack of personalization in traditional methods, this invention can dynamically adapt to different users' usage habits, permission levels, and operation intentions based on real-time user interaction data, providing users with personalized operation guidance tailored to their current needs. Other features and advantages of the embodiments of this application will be described in detail in the following detailed description section. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the following detailed description to explain the embodiments of this application, but do not constitute a limitation on the embodiments of this application. In the drawings: Figure 1 The illustration shows a flowchart of an operation guidance method according to an embodiment of this application; Figure 2 This schematic diagram illustrates a structural block diagram of an operation guidance device according to an embodiment of this application; Figure 3 The schematic diagram illustrates the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustration and explanation of the embodiments of this application and are not intended to limit the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0019] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0020] It should be noted that if the embodiments of this application involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.
[0021] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0022] Figure 1 The illustration shows a flowchart of an operation guidance method according to an embodiment of this application. Figure 1 As shown in the figure, this application provides an operation guidance method, which may include the following steps.
[0023] Step 101: Fuse the first question-and-answer information, the first behavior information, and the first permission information corresponding to the user terminal to obtain a fused feature vector.
[0024] In this embodiment, the first question-and-answer information includes the question text and answer text generated by the user through the user terminal. The user terminal integrates a large language model module, allowing the user to input question text (e.g., "How to generate financial statements") via natural language interaction. The large language model parses the question based on the system's knowledge base and returns a structured answer text (e.g., "To generate financial statements, click 'Financial Center → Report Management → Custom Report' in sequence"). The user terminal supports various electronic devices such as mobile phones, tablets, and computers, and has text input / output and data acquisition functions.
[0025] The first line of information includes page click data on the user's terminal. This page click data may include the sequence of pages the user logged into and the duration of time spent on each page.
[0026] The first permission information includes the user's permission level on the user terminal. For example, ordinary employees and administrators have different permission levels, which are pre-configured by the enterprise backend according to job roles, determining the scope of functions and operation permissions that users can access.
[0027] During the fusion process, the first question-and-answer information, the first action information, and the first permission information can be vectorized first, and then the three vectorized information can be fused to obtain the fused feature vector.
[0028] Step 102: Input the fused feature vector into the pre-acquired prediction model, and determine the probability of multiple functional items through the prediction model; In this embodiment of the application, the prediction model is a reinforcement learning model built based on a deep Q-network, and its network structure includes: Input layer: Used to receive fused feature vectors; Hidden layers: Multiple fully connected hidden layers are used. Each layer receives the output of the previous layer and undergoes a linear transformation through a weight matrix and a bias term. The Sigmoid activation function is used between layers to learn the non-linear relationships between features. Output layer: The number of neurons is the same as the total number of system functions. Each neuron corresponds to the Q-value of a function, which is the expected cumulative reward of that function under the conditions of the first question-and-answer information, the first action information, and the first permission information. After the model outputs the raw Q-values, they are normalized using the softmax function to convert the Q-values into a probability distribution, thus obtaining the probability of each function. For example, "generate financial statements" corresponds to a probability of 0.8, and "export customer data" corresponds to a probability of 0.15.
[0029] Among them, multiple functional items refer to various functional modules that can be operated by users in enterprise systems, such as "customer profile creation" and "order creation" in the sales process, "report generation" and "expense approval" in the financial process, and "role configuration" in the permission management process.
[0030] Step 103: Based on the order of probability from high to low among multiple functional items, select at least one target functional item from among the multiple functional items; In this embodiment, the top N most probable functional items can be selected as the target functional items, and the number of N can be customized by the enterprise according to its business scenario. During the screening process, the functional items with the highest probability are recommended first. For example, when the probability of "generating financial statements" is 0.8, the probability of "viewing report templates" is 0.1, and the probability of "printing reports" is 0.05, only the first two items are selected.
[0031] Step 104: Display at least one target function item on the user terminal's display interface.
[0032] In this embodiment, the target function can be displayed in real time through the "Smart Guidance" sidebar module on the right side of the user terminal interface for the user to select.
[0033] In this embodiment, by fusing the first question-and-answer information, first behavior information, and first permission information of the user terminal to generate a fused feature vector, the user's personalized needs and operation scenarios can be comprehensively captured. The fused feature vector is input into a prediction model to output the probability of function items and then filtered and displayed, changing the traditional method of generating operation guidance based on preset templates. Compared to the lack of personalization in traditional methods, this invention can dynamically adapt to different users' usage habits, permission levels, and operation intentions based on real-time user interaction data, providing users with personalized operation guidance tailored to their current needs.
[0034] In one embodiment of this application, the first question-and-answer information, the first behavior information, and the first permission information corresponding to the user terminal are fused to obtain a fused feature vector, including: Keyword extraction is performed on both the question text and the answer text, and the extracted keywords are converted into text feature vectors using a pre-trained word vector model; The page click data is arranged according to time series and processed by a recurrent neural network to obtain behavioral feature vectors; The permission levels are numerically encoded to obtain permission feature vectors; The text feature vector, behavior feature vector, and permission feature vector are concatenated to obtain the fused feature vector.
[0035] In this embodiment, the question text is the query content input by the user through natural language (e.g., "How to configure customer permissions"), and the answer text is the parsing result returned by the large language model based on the system knowledge base (e.g., "Configuring customer permissions requires operation under the path 'System Settings → User Management'"). Keyword extraction can use the TF-IDF algorithm to filter out common words such as "how" and "perform," retaining core business terms (e.g., "customer permissions," "system settings," and "user management"). The deduplicated keyword set can be input into a pre-trained word2vec model, which can map each keyword to a text feature vector of a specific dimension.
[0036] Page click data includes page click sequence and dwell time. Each page is assigned a unique integer ID, and the page click sequence is the sequence of IDs corresponding to the pages clicked by the user, arranged according to the click order. For example, if the click order is "Homepage → Customer List → Order Details → Payment Records", and the IDs for Homepage, Customer List, Order Details, and Payment Records are 1, 2, 3, and 4 respectively, then the corresponding page click sequence would be 1, 2, 3, 4. Dwell time is the duration spent on each page.
[0037] Subsequently, the page click sequence and dwell time are used as input to the RNN. The RNN encodes the page sequence into a one-hot vector. Finally, the standardized page click sequence vector and dwell time vector are concatenated in chronological order and input into the recurrent neural network for temporal feature extraction, outputting a fixed-length behavioral feature vector.
[0038] Role-based permission levels are pre-defined by the enterprise backend based on job attributes. For example, level 1 corresponds to ordinary employees, level 2 to sales supervisors, and level 3 to system administrators. Numerical encoding directly converts the level into a single-element integer vector, i.e., a permission feature vector.
[0039] During the feature fusion stage, the three types of vectors are concatenated according to their dimensions to form a complete fused feature vector, which is used to represent the user's state.
[0040] In this embodiment, by vectorizing and fusing the first question-and-answer information, the first behavior information, and the first permission information, a structured representation of multimodal data is achieved, which improves the accuracy of user state characterization and thus improves the accuracy of subsequent prediction of target function items.
[0041] In one embodiment of this application, after displaying at least one target function item on the display interface of a user terminal, the method further includes: When a user clicks on any target function item, the system records the second question and answer information, the second behavior information, and the second permission information corresponding to that click. At preset time intervals, the prediction model is updated using the second question-and-answer information, the second behavior information, and the second permission information.
[0042] In this embodiment, the second question-and-answer information includes the text of the last question before the user clicked the function item and the answer text of the large language model. The second behavior information includes the user's page click data on the current page. The second permission information is consistent with the permission level of the user's currently logged-in account.
[0043] The preset time period is set according to the enterprise's business cycle, preferably 7 days. Within each update cycle, all click operation data (i.e., second question and answer, second behavior, and second permission information) collected within the cycle are converted into training samples and stored in the experience replay buffer for periodically updating and training the prediction model. For the specific training process, please refer to the following embodiment, which will not be elaborated here.
[0044] In this embodiment, by recording user click feedback in real time and periodically updating the prediction model, the prediction accuracy of the prediction model can be further improved.
[0045] In one embodiment of this application, the prediction model is a model trained based on a deep Q-network, and the training process includes: Construct an experience replay buffer, which includes multiple experience data, each of which includes the user's corresponding current state, action, reward value, and next state. The first-depth Q-network was trained using multiple empirical datasets to obtain the predicted Q-values; The second deep Q-network was trained using multiple empirical data sets to obtain the target Q-value. The second deep Q-network and the first deep Q-network have the same network structure but different network parameters. The loss function value is determined based on the predicted Q value and the target Q value; If the loss function value does not meet the preset training stopping condition, return to the previous step and train the first deep Q network using multiple empirical data to obtain the predicted Q value until the loss function value meets the preset training stopping condition, at which point the first deep Q network is determined as the prediction model.
[0046] In this embodiment, the current state is a historical fusion feature vector generated by the method described in the above embodiments, representing the comprehensive state of the user's question content, operation trajectory, and permission level; the action is the recommended function item corresponding to the current state; the reward value is the incentive value fed back by the user after performing the action according to preset rules, and the specific calculation method is described in the following embodiments; the next state is the new fusion feature vector after the user performs the action. It should be noted that the multiple experience data in this embodiment can be multiple historical experience data.
[0047] Before training a deep Q-network, a batch of empirical data can be randomly selected from multiple empirical datasets as training samples to break the temporal correlation between data and improve the model's generalization ability.
[0048] The network structures of the first and second deep Q-networks are the same as those of the prediction model in the above embodiments, and will not be repeated here. The parameter updates of the second deep Q-network lag behind those of the first deep Q-network. Specifically, the model parameters of the first deep Q-network are copied to the second deep Q-network at preset iteration intervals. The preferred number of iterations is 100.
[0049] In calculating the predicted Q-value and the target Q-value, the input data to the deep Q-network can be the current state, and the output is calculated using the action, reward value, and next state.
[0050] Based on the predicted Q-value and the target Q-value, the loss function value is determined. The root mean square error (RMSE) can be used as the loss function, and its expression is:
[0051] Where L is the loss function value, N It is the amount of empirical data. It is the i-th predicted Q value. It is the i-th objective Q value.
[0052] Preset training stopping conditions may include the loss function converging or reaching the maximum number of iterations.
[0053] In this embodiment, by constructing an experience replay buffer to store multi-dimensional experience data, and combining it with a dual-depth Q-network architecture, the stability and convergence speed of model training are improved by utilizing the root mean square error loss function and iterative training mechanism. This, in turn, enhances the model's prediction accuracy.
[0054] In one embodiment of this application, the reward value is determined through the following process: Based on preset evaluation indicators and recommended function items corresponding to actions, the interaction value of users to actions is calculated. The recommended function items are the function items predicted by the prediction model based on the current state. The reward value is determined based on the interaction value; The evaluation metrics include at least one of the following: click behavior reward, feature completion reward, dwell time reward, and penalty.
[0055] In this embodiment, an action refers to the user's actual click on a recommended function item on the display interface. The recommended function item is a function item predicted by the prediction model based on the current state in empirical data. The interaction value can be determined by identifying the difference between the action and the recommended function item.
[0056] Specifically, click behavior rewards refer to the positive reward given when a user clicks on a function recommended by the prediction model. This reward can be fine-tuned based on the importance of the function and the frequency of user clicks. For example, a higher reward can be given when the user's action corresponds to a core function or a frequently used recommended function.
[0057] The formula is:
[0058] in, Let α be the first interaction value corresponding to the click behavior reward, α be the first adjustment coefficient, Importance(f) be the importance score of action f, and Frequency(f) be the usage frequency of action f.
[0059] Stay time reward: The time a user spends on a particular feature page is also an important metric for measuring whether that feature meets user needs. Therefore, rewards can be given based on the time users spend on a page. To encourage users to stay on useful pages for longer periods, the dwell time reward can be set as an increasing function, such as a logarithmic or exponential function. Therefore, when the action indicates staying on a certain page, the second interaction value corresponding to the dwell time reward can be determined using the following formula.
[0060] The formula is:
[0061] in, It is the second interaction value corresponding to the dwell time reward, and β is the second adjustment coefficient used to adjust the sensitivity of the reward.
[0062] Feature completion reward: In some cases, users may need to complete a series of steps or tasks to use a feature. Feature completion rewards refer to incentives given to users based on the degree of feature completion to encourage them to complete these steps. This reward can be fixed or fine-tuned based on the completion level. Therefore, when an action represents the completion level of a feature, the third interaction value corresponding to the feature completion reward can be determined using the following formula.
[0063] formula:
[0064] in, The third interaction value corresponding to the feature completion reward is δ, which is the third adjustment coefficient, and CompletionRate is the feature completion rate (e.g., the proportion of steps completed).
[0065] Penalties: In addition to positive rewards, penalties can be designed to prevent the model from generating recommendations that do not meet user needs. For example, when a user clicks on a system-recommended feature but then quickly returns or closes the page, a penalty can be imposed.
[0066] formula:
[0067] in, ϵ is the fourth interaction value corresponding to the penalty item, ϵ is the fourth adjustment coefficient, and BounceRate is the user's bounce rate (i.e., the proportion of users who leave quickly after clicking).
[0068] The interaction values corresponding to each evaluation indicator are weighted and summed according to preset weights to form the final reward value, which is determined by the following formula:
[0069] in, For the reward value, w1 is the first weight coefficient corresponding to the click behavior reward, w2 is the second weight coefficient corresponding to the function completion reward, w3 is the third weight coefficient corresponding to the dwell time reward, and w4 is the fourth weight coefficient corresponding to the penalty item. For the explanation of the other parameters, please refer to the above embodiment, which will not be repeated here.
[0070] In this embodiment, model training is performed by integrating click behavior, dwell time, function completion rate, and penalty terms to avoid the bias of a single indicator in the model training process. This improves the comprehensiveness of the model and thus enhances its accuracy.
[0071] Figure 2 A schematic diagram of the operation guidance device provided in another embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.
[0072] Reference Figure 2 The operation guidance device 200 may include: The fusion module 201 is used to fuse the first question-and-answer information, the first behavior information, and the first permission information corresponding to the user terminal to obtain a fused feature vector. The first question-and-answer information includes the question text and answer text generated by the user through the user terminal; the first behavior information includes the page click data of the user on the user terminal; and the first permission information includes the permission level of the user on the user terminal. The output module 202 is used to input the fused feature vector into the pre-acquired prediction model, and to determine the probability of multiple functional items through the prediction model; The filtering module 203 is used to filter out at least one target function item from multiple function items based on the order of their probabilities from high to low. Display module 204 is used to display at least one target function item on the display interface of the user terminal.
[0073] Optionally, the fusion module 201 includes: The extraction submodule is used to extract keywords from the question text and the answer text respectively, and to convert the extracted keywords into text feature vectors using a pre-trained word vector model; The sorting submodule is used to sort page click data according to time series and process it through a recurrent neural network to obtain behavioral feature vectors; The encoding submodule is used to numerically encode the permission levels to obtain permission feature vectors. The concatenation submodule is used to concatenate text feature vectors, behavior feature vectors, and permission feature vectors to obtain a fused feature vector.
[0074] Optionally, the operation guidance device 200 further includes: The recording module is used to record the second question and answer information, the second behavior information, and the second permission information corresponding to the click operation when a user clicks on any target function item. The update module is used to update the prediction model at preset time intervals using second question and answer information, second behavior information, and second permission information.
[0075] Optionally, the operation guidance device 200 further includes: The building module is used to build the experience replay buffer, which includes multiple experience data, each of which includes the user's corresponding current state, action, reward value, and next state; The first training module is used to train the first deep Q network using multiple empirical data to obtain the predicted Q value; The second training module is used to train the second deep Q network using multiple empirical data to obtain the target Q value. The second deep Q network has the same network structure as the first deep Q network, but the network parameters are different. The first determining module is used to determine the loss function value based on the predicted Q value and the target Q value; The second determination module is used to return to training the first deep Q network using multiple empirical data to obtain the predicted Q value when the loss function value does not meet the preset training stopping condition, and then determine the first deep Q network as the prediction model.
[0076] Optionally, the operation guidance device 200 is specifically used for: At preset intervals, the model parameters of the first deep Q-network are copied to the second deep Q-network.
[0077] Optionally, the operation guidance device 200 further includes: The calculation module is used to calculate the user's interaction value with an action based on preset evaluation indicators and recommended function items corresponding to the action. The recommended function items are the function items predicted by the prediction model based on the current state. The third determination module is used to determine the reward value based on the interaction value; The evaluation metrics include at least one of the following: click behavior reward, feature completion reward, dwell time reward, and penalty.
[0078] Optionally, the loss function value is calculated using the following formula:
[0079] Where L is the loss function value and N is the number of empirical data points. It is the i-th predicted Q value. It is the i-th objective Q value.
[0080] Figure 3 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0081] The device may include a processor 301 and a memory 302 storing program instructions.
[0082] When processor 301 executes the program, it implements the steps in any of the above method embodiments.
[0083] For example, the program can be divided into one or more modules / units, one or more of which are stored in memory 302 and executed by processor 301 to complete this application. The one or more modules / units can be a series of program instruction segments capable of performing a specific function, which describe the execution process of the program in the device.
[0084] Specifically, the processor 301 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0085] Memory 302 may include mass storage for data or instructions. For example, and not limitingly, memory 302 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 302 may include removable or non-removable (or fixed) media. Where appropriate, memory 302 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 302 is non-volatile solid-state memory.
[0086] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.
[0087] The processor 301 implements any of the methods described in the above embodiments by reading and executing program instructions stored in the memory 302.
[0088] In one example, the electronic device may also include a communication interface 303 and a bus 310. The processor 301, memory 302, and communication interface 303 are connected via the bus 310 and communicate with each other.
[0089] The communication interface 303 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0090] Bus 310 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 310 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0091] Furthermore, in conjunction with the methods in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores program instructions; when these program instructions are executed by a processor, they implement any of the methods in the above embodiments.
[0092] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0093] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0094] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.
[0095] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0096] The functional modules shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on machine-readable media or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable media" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer grids such as the Internet, intranets, etc.
[0097] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0098] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0099] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. An operation guidance method, characterized in that, The method includes: The first question-and-answer information, the first behavior information, and the first permission information corresponding to the user terminal are fused to obtain a fused feature vector. The first question-and-answer information includes the question text and answer text generated by the user through the user terminal; the first behavior information includes the page click data of the user on the user terminal; and the first permission information includes the user's permission level on the user terminal. The fused feature vector is input into a pre-acquired prediction model, and the probability of multiple functional items is determined by the prediction model. Based on the probability of the plurality of functional items in descending order, at least one target functional item is selected from the plurality of functional items. The at least one target function item is displayed on the user terminal's display interface.
2. The method according to claim 1, characterized in that, The process of fusing the first question-and-answer information, the first behavior information, and the first permission information corresponding to the user terminal to obtain a fused feature vector includes: Keyword extraction is performed on the question text and the answer text respectively, and the extracted keywords are converted into text feature vectors using a pre-trained word vector model; The page click data is arranged according to a time series and processed by a recurrent neural network to obtain a behavioral feature vector; The permission level is numerically encoded to obtain a permission feature vector; The text feature vector, the behavior feature vector, and the permission feature vector are concatenated to obtain the fused feature vector.
3. The method according to claim 1, characterized in that, After displaying the at least one target function item on the user terminal's display interface, the method further includes: When a user clicks on any target function item, the system records the second question and answer information, the second behavior information, and the second permission information corresponding to that click. At preset time intervals, the prediction model is updated using the second question-and-answer information, the second behavior information, and the second permission information.
4. The method according to claim 1, characterized in that, The prediction model is a model trained based on a deep Q-network, and the training process includes: An experience replay buffer is constructed, which includes multiple experience data, each of which includes the user's corresponding current state, action, reward value, and next state; The first deep Q-network is trained using the aforementioned empirical data to obtain the predicted Q-value; The second deep Q-network is trained using the aforementioned empirical data to obtain the target Q-value. The second deep Q-network and the first deep Q-network have the same network structure but different network parameters. Based on the predicted Q-value and the target Q-value, the loss function value is determined; If the loss function value does not meet the preset training stopping condition, the process returns to training the first deep Q network using the multiple empirical data to obtain the predicted Q value, until the loss function value meets the preset training stopping condition, at which point the first deep Q network is determined as the prediction model.
5. The method according to claim 4, characterized in that, After determining the loss function based on the predicted Q-value and the target Q-value, the method further includes: At preset intervals, the model parameters of the first deep Q network are copied to the second deep Q network.
6. The method according to claim 4, characterized in that, The reward value is determined through the following process: Based on preset evaluation indicators and recommended function items corresponding to the action, the interaction value of the user to the action is calculated, and the recommended function items are the function items predicted by the prediction model based on the current state. The reward value is determined based on the interaction value. The evaluation metrics include at least one of the following: click behavior reward, function completion reward, dwell time reward, and penalty item.
7. The method according to claim 4, characterized in that, The loss function value is calculated using the following formula: Where L is the loss function value, N It is the amount of empirical data. It is the i-th predicted Q value. It is the i-th objective Q value.
8. An operation guidance device, characterized in that, The device includes: The fusion module is used to fuse the first question-and-answer information, the first behavior information, and the first permission information corresponding to the user terminal to obtain a fused feature vector. The first question-and-answer information includes the question text and answer text generated by the user through the user terminal; the first behavior information includes the page click data of the user on the user terminal; and the first permission information includes the user's permission level on the user terminal. The output module is used to input the fused feature vector into the pre-acquired prediction model, and to determine the probability of multiple functional items through the prediction model; A filtering module is used to filter out at least one target function item from the plurality of function items based on the order of their probabilities from high to low. The display module is used to display the at least one target function item on the display interface of the user terminal.
9. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the operation guidance method as described in any one of claims 1-7.
10. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores instructions for causing the machine to perform the operation guidance method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Individual application function recommendation method with multi-mode clues and system thereof
CN106227815A
Personalized display method and device
CN111562963A
Personalized menu recommendation method and system, computer equipment and storage medium
CN116737039A