Recommendation method and device
By establishing and evaluating the trajectory tree of the policy execution path, the problem of low accuracy of manual processing strategies in the prior art is solved, and automated policy determination and recommendation are realized, improving the accuracy and user experience of the policy.
Patent Information
- Application Number
- CN202510237986.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, the formulation of product processing strategies requires relying on manual experience, resulting in low accuracy of the strategy.
By responding to user's information query request, preliminary information of the target product is obtained, and the first trajectory tree is established based on the preliminary information and policy set. The evaluation model is used to determine the target candidate policies in each level, and ultimately recommend the policy execution path including the target candidate policies to the user.
It realizes automatic determination of the policy execution path that matches the target product, improves the accuracy of processing policies, and improves the user experience.
Smart Images

Figure CN120179894A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a recommendation method and apparatus. Background Art
[0002] In order to ensure efficient processing of products (such as sales, transportation, development, etc.), it is generally necessary to pre-formulate a processing strategy corresponding to the product. Currently, the processing strategy of the product can be formulated manually, but this method relies on manual experience and the accuracy of the formulated processing strategy is low. Summary of the Invention
[0003] An embodiment of this application provides a recommendation method, including: in response to obtaining a user's information query request for a target product, obtaining preliminary information of the target product; establishing a first trajectory tree according to the preliminary information and a policy set, the policy set includes multiple initial policies, the first trajectory tree includes multiple levels that are executed sequentially, each level includes at least one candidate policy as a node, and the candidate policy is at least determined according to the preliminary information and historical information corresponding to the initial policy; determining a target candidate policy in each level according to the evaluation results corresponding to the nodes in the first trajectory tree, the evaluation results are determined after evaluating each candidate policy by invoking a first evaluation model; recommending to the user at least one first policy execution path including the target candidate policy.
[0004] In some embodiments, the first evaluation model belongs to a target intelligent agent, and the target intelligent agent further includes a first target large model, and establishing the first trajectory tree according to the preliminary information and the policy set is executed based on the first target large model;
[0005] Alternatively, the target intelligent agent further includes a second target large model and a second evaluation model, the second target large model is used to generate a second trajectory tree according to the information query request and the preliminary information; the second evaluation model is used to evaluate candidate policies in multiple levels of the second trajectory tree to determine target candidate policies, so as to obtain a second policy execution path.
[0006] In some embodiments, the first evaluation model belongs to a target intelligent agent, and the target intelligent agent further includes a third target large model and a third evaluation model, the third target large model is used to generate new candidate policies according to interaction data associated with the target candidate policies, and the new candidate policies are different from the initial policies; the third evaluation model is used to evaluate the new candidate policies;
[0007] Or,
[0008] In the case where the first policy execution path does not meet the requirements, the third target large model is used to generate a third trajectory tree according to the interaction data associated with the historical policy execution path. Each candidate policy in the third trajectory tree is a new candidate policy or different from the candidate policy in the hierarchy. The new candidate policy is different from the initial policy, and the historical policy execution path includes at least the first policy execution path. The third evaluation model is used to evaluate each candidate policy in the third trajectory tree.
[0009] In some embodiments, after recommending at least one first policy execution path including the target candidate policy to the user, it further includes: obtaining feedback information of the user on the first policy execution path; in the case where the feedback information meets the target condition, determining the current target candidate policy in the first policy execution path, and adjusting the first policy execution path according to the execution situation of the candidate policies in the current hierarchy where the current target candidate policy is located. The target condition includes one of the following: not executing the current target candidate policy; the current target candidate policy fails to execute.
[0010] In some embodiments, the adjusting the first policy execution path according to the execution situation of the candidate policies in the current hierarchy where the current target candidate policy is located includes: determining the number of executable candidate policies in the current hierarchy; in the case where the number is zero, determining the layer type of the current hierarchy, and the layer type includes the hierarchy where the root node is located or the hierarchy where the non-root node is located; according to the layer type, adjusting the first policy execution path based on the third target large model.
[0011] In some embodiments, the adjusting the first policy execution path based on the third target large model according to the layer type includes: in the case where the layer type is the hierarchy where the root node is located, generating the third trajectory tree based on the third target large model according to the interaction data associated with the historical policy execution path; evaluating each candidate policy in the third trajectory tree based on the third evaluation model, determining the target candidate policy in the third trajectory tree to obtain the third policy execution path; using the third policy execution path to replace the first policy execution path; in the case where the layer type is the hierarchy where the non-root node is located, generating a plurality of the new candidate policies based on the third target large model according to the interaction data associated with the target candidate policy in the first trajectory tree; evaluating each of the new candidate policies based on the third evaluation model, determining the target new candidate policy, and using the target new candidate policy to replace the current target candidate policy.
[0012] In some embodiments, after replacing the current target candidate policy with the target new candidate policy, the method further includes: adjusting parameters of the third target large model according to the execution situation of the target new candidate policy.
[0013] In some embodiments, after determining the number of executable candidate policies in the current layer, the method further includes: in the case where the number is multiple, selecting a target executable candidate policy from the executable candidate policies, and replacing the current target candidate policy with the target executable candidate policy; in the case where the number is one, replacing the current target candidate policy with the executable candidate policy.
[0014] In some embodiments, before obtaining the preliminary information of the target product, the method further includes: in response to obtaining the information query request, determining the request type of the information query request; in the case where the request type is a factual request, querying a target knowledge graph according to the information query request to determine first information, the target knowledge graph belonging to a target intelligent agent; recommending the first information to the user; in the case where the request type is a non-factual request, determining second information according to the information query request by using a fourth target model, the fourth target model belonging to the target intelligent agent; recommending the second information to the user; in the case where the second information cannot be determined by using the fourth target model, performing the step of obtaining the preliminary information of the target product.
[0015] An embodiment of the present application further provides a recommendation device, including: an obtaining module, configured to obtain preliminary information of a target product in response to obtaining an information query request of a user for the target product; a building module, configured to build a first trajectory tree according to the preliminary information and a policy set, the policy set including a plurality of initial policies, the first trajectory tree including a plurality of sequentially executed layers, each layer including at least one candidate policy as a node, the candidate policy being determined at least according to the preliminary information and historical information corresponding to the initial policy; a determining module, configured to determine a target candidate policy in each layer according to evaluation results corresponding to the nodes in the first trajectory tree, the evaluation results being determined by calling a first evaluation model to evaluate the candidate policies; and a recommending module, configured to recommend at least one first policy execution path including the target candidate policy to the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 Flow of the recommendation method according to the embodiment of the present application Figure 1 ;
[0018] Figure 2 Flow of the recommendation method according to the embodiment of the present application Figure 2 ;
[0019] Figure 3 Flow of adjusting the execution path of the first strategy according to the embodiment of the present application Figure 1 ;
[0020] Figure 4 Flow of adjusting the execution path of the first strategy according to the embodiment of the present application Figure 2 ;
[0021] Figure 5 Schematic diagram of the first trajectory tree according to the embodiment of the present application Figure 1 ;
[0022] Figure 6 Schematic diagram of the first trajectory tree according to the embodiment of the present application Figure 2 ;
[0023] Figure 7 Schematic diagram of the first trajectory tree according to the embodiment of the present application Figure 3 ;
[0024] Figure 8 Schematic diagram of the first trajectory tree according to the embodiment of the present application Figure 4 ;
[0025] Figure 9 Schematic diagram of the second or third trajectory tree according to the embodiment of the present application;
[0026] Figure 10 Schematic diagram of the first trajectory tree according to the embodiment of the present application Figure 5 ;
[0027] Figure 11 Schematic diagram of the first trajectory tree according to the embodiment of the present application Figure 6 ;
[0028] Figure 12 Structural block diagram of the recommendation device according to the embodiment of the present application. Detailed implementation manners
[0029] Various solutions and features of the present application are described herein with reference to the accompanying drawings.
[0030] It should be understood that various modifications can be made to the embodiments applied herein. Therefore, the above description should not be regarded as a limitation, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of the present application.
[0031] The drawings included in and forming a part of the specification illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, are used to explain the principles of the present application.
[0032] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting examples with reference to the drawings.
[0033] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application.
[0034] When combined with the drawings, the above and other aspects, features and advantages of the present application will become more apparent in view of the following detailed description.
[0035] Specific embodiments of the present application will be described hereinafter with reference to the drawings; however, it should be understood that the embodiments claimed are merely examples of the present application and can be implemented in various ways. Well-known and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but are merely used as a basis and representative basis for the claims to teach those skilled in the art to use the present application in substantially any suitable detailed structure in a variety of ways.
[0036] This specification may use the phrases "in one embodiment", "in another embodiment", "in yet another embodiment" or "in other embodiments", which may each refer to one or more of the same or different embodiments according to the present application.
[0037] A recommended method for an embodiment of the present application, after obtaining a user's information query request for a target product, establishes a first trajectory tree according to the preliminary information of the target product and a set of strategies, determines the target candidate strategies in each level of the first trajectory tree according to the evaluation results corresponding to the nodes in the first trajectory tree, and then recommends at least one first strategy execution path including the target candidate strategies to the user. In this way, according to the user's information query request, a strategy execution path matching the target product is automatically determined and recommended to the user, thereby realizing the automatic determination of the product processing strategy and recommending it to the user, improving the accuracy of the processing strategy and enhancing the user experience.
[0038] As Figure 1 shown, the recommended method includes the following steps:
[0039] Step S101, in response to obtaining a user's information query request for a target product, obtain the preliminary information of the target product.
[0040] In this embodiment, the information query request may include, for example, a query request for any one of information such as product price inquiry, product transportation route planning, and product development strategy. The information query request can be obtained according to the input information of the user on the human-computer interaction interface, or according to the voice issued by the user. The preliminary information can be generated in real time using a preliminary information evaluation model corresponding to the target product, or can be information generated and stored in advance using the preliminary information evaluation model, or can also be information input by the user according to expert experience.
[0041] For example, if the information query request is a product price inquiry, the preliminary information can be the basic price of the target product determined in advance using a price evaluation model. If the information query request is product transportation route planning, the preliminary information can be the preliminary transportation route determined in advance according to a route planning model or expert experience.
[0042] Step S102: Establish a first trajectory tree according to the preliminary information and the strategy set. The strategy set includes a plurality of initial strategies. The first trajectory tree includes multiple levels that are executed sequentially. Each level includes at least one candidate strategy serving as a node, and the candidate strategy is determined at least according to the preliminary information and the historical information corresponding to the initial strategy.
[0043] In this embodiment, the strategy set is a pre-set set corresponding to the target product and including a plurality of initial strategies. After obtaining the preliminary information, a plurality of candidate strategies are determined according to the preliminary information and the historical information corresponding to the strategy set, and a first trajectory tree is established according to the candidate strategies. The first trajectory tree includes multiple levels that are executed sequentially, and each level includes at least one candidate strategy serving as a node. Among them, the historical information corresponding to the strategy set includes, for example, the user's historical decisions or historical feedback on the initial strategies, the historical sales situation or historical profit situation of the products corresponding to the initial strategies, etc. By considering the historical information corresponding to the strategy set when determining the candidate strategies, the adoption of inappropriate initial strategies can be avoided, making the candidate strategies more in line with the actual situation of the target product and improving the accuracy of the candidate strategies.
[0044] For example, when the information query request is a product price inquiry and the preliminary information is the base price, each initial strategy in the strategy set may include reducing the base price, increasing the base price, suggesting an increase in quantity, suggesting a decrease in quantity, suggesting a replacement with a similar product, upgrading product accessories (such as CPU upgrade), downgrading product accessories, and the price and quantity of peripheral products. The candidate strategies can be reducing the base price (including one or more price reduction ratios), increasing the base price (including one or more price increase ratios), suggesting an increase in quantity (including one or more specific quantities), suggesting a decrease in quantity (including one or more specific quantities), suggesting a replacement with a similar product (including the names of one or more similar products), upgrading product accessories (including one or more product accessories for upgrading), downgrading product accessories (including one or more product accessories for downgrading), and the price and quantity of peripheral products (including one or more peripheral products and the corresponding prices and quantities).
[0045] Figure 5 Schematic of the first trajectory tree of the embodiment of the present application Figure 1 , Figure 5 In this example, the target product is a laptop computer. The first trajectory tree includes three levels executed sequentially from top to bottom. The candidate strategies corresponding to the nodes in the first level include, for example, a price decrease of x1%, a quantity increase of x2%; a price increase of y1%, a quantity decrease of y2%; a price increase of z1%; an upward upgrade of the CPU product sales price and quantity; a downward downgrade of the CPU product sales price and quantity, etc. The candidate strategies corresponding to the nodes in the second level include, for example, the price and quantity of other similar products; the sales price and quantity of peripheral products, such as a computer stand; a product price decrease and a quantity increase, etc. The candidate strategies corresponding to the nodes in the third level include, for example, the sales price and quantity of peripheral products; the price and quantity of other similar products, etc.
[0046] Step S103: Determine the target candidate strategy in each level according to the evaluation results corresponding to the nodes in the first trajectory tree, where the evaluation results are determined by calling a first evaluation model to evaluate each candidate strategy.
[0047] In this embodiment, each candidate strategy is evaluated one or more times (such as cross-state voting) by calling a first evaluation model to determine the evaluation results corresponding to the nodes in the first trajectory tree. The evaluation results can represent the likelihood of successfully executing the candidate strategy. According to the evaluation results, the target candidate strategy in each level is determined. For example, one or more candidate strategies with the highest likelihood of successful execution in each level are determined as the target candidate strategy.
[0048] Step S104: Recommend at least one first strategy execution path including the target candidate strategy to the user.
[0049] In this embodiment, the target candidate policies in each level can form at least one first policy execution path, and at least one first policy execution path is recommended to the user. For example, in the first trajectory tree, at least one first policy execution path is displayed according to the target display effect, so as to recommend the first policy execution path to the user. Or a window including at least one first policy execution path can also be popped up, so as to recommend the first policy execution path to the user. Taking Figure 5 the first trajectory tree in [reference number] as an example, the first policy execution path may include, for example, successively performing a price reduction of x1% and a quantity increase of x2% on the basis of the base price; the prices and quantities of other products in the same category; the selling prices and quantities of peripheral products.
[0050] In some embodiments of the present application, each node in the first trajectory tree further includes explanatory information about the candidate policy, such as the estimated conversion rate, the estimated impact on the conversion rate and profit of subsequent nodes, etc., so as to provide more reference information and improve the user experience while recommending the policy to the user. For example, as Figure 6 shown in [reference number], if a price reduction of x1% and a quantity increase of x2% are adopted, the estimated conversion rate is 95%, and the conversion rate is high. If further recommended according to the prices and quantities of other products in the same category, the conversion rate increases by k%, and the profit reaches m%.
[0051] In some embodiments of the present application, the candidate policy corresponding to the edit instruction can also be adjusted in response to the user's edit instruction for the first trajectory tree. For example, the user can adjust the price reduction or increase ratio, the prices and quantities of other products in the same category, etc., so that the user can adjust the first trajectory tree according to needs, improving the user experience.
[0052] The recommendation method according to the embodiment of the present application obtains the preliminary information of the target product in response to obtaining the user's information query request for the target product; establishes a first trajectory tree according to the preliminary information and a policy set, the policy set includes a plurality of initial policies, the first trajectory tree includes multiple levels executed in sequence, each level includes at least one candidate policy as a node, and the candidate policy is determined at least according to the preliminary information and the historical information corresponding to the initial policy; determines the target candidate policy in each level according to the evaluation results corresponding to each node in the first trajectory tree, and the evaluation results are determined by calling a first evaluation model to evaluate each candidate policy; recommends at least one first policy execution path including the target candidate policy to the user. In this way, according to the user's information query request, the policy execution path matching the target product is automatically determined and recommended to the user, so as to automatically determine the product processing policy and recommend it to the user, improving the accuracy of the processing policy and enhancing the user experience.
[0053] In some embodiments of the present application, the first evaluation model belongs to the target agent, and the target agent further includes a first target large model. The establishment of the first trajectory tree based on the preliminary information and the policy set is executed based on the first target large model;
[0054] Alternatively, the target agent further includes a second target large model and a second evaluation model.
[0055] The second target large model is used to generate a second trajectory tree according to the information query request and the preliminary information;
[0056] The second evaluation model is used to evaluate the candidate policies at multiple levels in the second trajectory tree to determine the target candidate policy, so as to obtain a second policy execution path.
[0057] In this embodiment, the target agent is an agent that can perceive the environment and take actions to achieve specific goals. It can be software, hardware, or a system, and has autonomy, adaptability, and interaction capabilities. The target agent includes a first evaluation model and a first target large model. By calling the first target large model from the target agent, a first trajectory tree is established according to the preliminary information and the policy set, and then the first evaluation model is called from the agent to evaluate each candidate policy in the first trajectory tree. Among them, the first target large model can be established based on, for example, a thought tree.
[0058] Alternatively, the target agent further includes a second target large model and a second evaluation model. In the case where the first policy execution path does not meet the requirements, such as the user does not accept the first policy execution path, or the first policy execution path fails, the second target large model can be called from the target agent to generate a second trajectory tree according to the information query request and the preliminary information, and then the second evaluation model is called from the target agent to evaluate each candidate policy in the second trajectory tree to determine the target candidate policy, so as to obtain a second policy execution path. Among them, the second target large model is different from the first target large model. The second target large model can be established based on, for example, a thought chain, so as to ensure that the second policy execution path is different from the first policy execution path, and improve the possibility of successfully executing the policy. Optionally, the first evaluation model and the second evaluation model can be the same or different.
[0059] In some embodiments of the present application, as an alternative, after obtaining the preliminary information of the target product, the second target large model can also be directly called from the target agent to generate a second trajectory tree according to the information query request and the preliminary information; the second evaluation model is called from the target agent to evaluate each candidate policy in the second trajectory tree, determine the target candidate policy in the second trajectory tree, so as to obtain a second policy execution path; and the second policy execution path is recommended to the user.
[0060] By calling the first target large model and the first evaluation model from the target agent to determine the first policy execution path, more accurate and efficient determination of the first policy execution path is achieved. By calling the second target large model and the second evaluation model from the target agent to determine the second policy execution path, more accurate and efficient determination of the second policy execution path is achieved.
[0061] In some embodiments of the present application, the first evaluation model belongs to the target agent, and the target agent further includes a third target large model and a third evaluation model.
[0062] The third target large model is used to generate a new candidate policy according to the interaction data associated with the target candidate policy, and the new candidate policy is different from the initial policy.
[0063] The third evaluation model is used to evaluate the new candidate policy.
[0064] Or,
[0065] In the case where the first policy execution path does not meet the requirements,
[0066] The third target large model is used to generate a third trajectory tree according to the interaction data associated with the historical policy execution path. Each candidate policy in the third trajectory tree is a new candidate policy or different from the candidate policy in the hierarchy; the new candidate policy is different from the initial policy, and the historical policy execution path at least includes the first policy execution path.
[0067] The third evaluation model is used to evaluate each candidate policy in the third trajectory tree.
[0068] In this embodiment, the target agent further includes a third target large model and a third evaluation model. When all candidate policies at the level where the target candidate policy is located do not meet the requirements, the third target large model can be called to generate a new candidate policy according to the interaction data associated with the target candidate policy. The new candidate policy is different from the initial policy. Then, the third evaluation model is called to evaluate the generated multiple new candidate policies to determine the target new candidate policy recommended to the user, thereby ensuring reliable policy recommendation to the user.
[0069] Or, in the case where the first policy execution path does not meet the requirements, the third target large model can be called from the target agent to generate a third trajectory tree according to the interaction data associated with the historical policy execution path. Each candidate policy in the third trajectory tree is a new candidate policy or different from the candidate policy in the hierarchy, and the new candidate policy is different from the initial policy, thereby avoiding recommending policies that do not meet the requirements to the user again. Then, the third evaluation model is called to evaluate each candidate policy in the third trajectory tree to determine the target candidate policy, so as to obtain the third policy execution path.
[0070] The third target large model is different from the first target large model. The third target large model can be established based on, for example, the deep thinking model, so as to ensure that the third policy execution path is different from the first policy execution path, improving the possibility of successfully executing the policy. Optionally, the first evaluation model and the third evaluation model can be the same or different.
[0071] In some embodiments of the present application, after recommending at least one first policy execution path including the target candidate policy to the user, as Figure 2 shown, the following steps are further included:
[0072] Step S105, obtaining feedback information of the user on the first policy execution path.
[0073] In this embodiment, after recommending the first policy execution path to the user, the feedback information of the user is monitored. The feedback information can include, for example, the successful execution of the target candidate policy, the failure of the target candidate policy to execute, not executing the target candidate policy, etc. In some embodiments, after recommending the first policy execution path to the user, a selection interface for obtaining feedback information is displayed to the user. Corresponding options can be displayed in the selection interface. For example, an option of whether to adopt the target candidate policy and an option of whether the target candidate policy is successfully executed are displayed in the selection interface, so as to guide the user to provide feedback information and achieve more efficient acquisition of feedback information.
[0074] Step S106, when the feedback information meets the target condition, determining the current target candidate policy in the first policy execution path, and adjusting the first policy execution path according to the execution situation of the candidate policies in the current layer where the current target candidate policy is located.
[0075] In this embodiment, the target condition includes one of the following: not executing the current target candidate policy; the failure of the current target candidate policy to execute. If the feedback information meets the target condition, it indicates that the current target candidate policy is not appropriate and other policies need to be recommended. According to the execution situation of the candidate policies in the current layer where the current target candidate policy is located, the first policy execution path is adjusted. Subsequently, the adjusted first policy execution path can be recommended to the user, thereby realizing the automatic adjustment of the first policy execution path.
[0076] In some embodiments of the present application, adjusting the first policy execution path according to the execution situation of the candidate policies in the current layer where the current target candidate policy is located, as Figure 3 shown, includes the following steps:
[0077] Step S1061, determining the number of executable candidate policies in the current layer.
[0078] In this embodiment, if only the current target candidate policy is included in the current layer, the number of executable candidate policies in the current layer is zero. If the current layer includes the current target candidate policy and other candidate policies, the number of executable candidate policies is determined based on the other candidate policies. In some embodiments, the number of times of receiving the target feedback information corresponding to the current layer can be determined, and the target feedback information meets the target condition; when the number of times reaches the target number of times or all candidate policies in the current layer have been recommended, the number is determined to be zero. When the number of times does not reach the target number of times and there are candidate policies in the current layer that have not been recommended, the number is determined not to be zero.
[0079] Step S1062, when the number is zero, determine the layer type of the current layer, where the layer type includes the layer where the root node is located or the layer where a non-root node is located.
[0080] In this embodiment, when the number of executable candidate policies in the current layer is zero, it means that all candidate policies in the current layer are not suitable. Determine the layer type of the current layer, where the layer type includes the layer where the root node is located or the layer where a non-root node is located. For example, as Figure 5 shown, the layer where the price drops by x1% and the quantity increases by x2% corresponds to the layer where the root node is located, and the layer where the price and quantity of other products of the same type correspond to the layer where a non-root node is located.
[0081] Step S1063, according to the layer type, adjust the first policy execution path based on the third target large model.
[0082] In this embodiment, according to the layer type, the third target large model is called from the target agent to adjust the first policy execution path.
[0083] By determining the number of executable candidate policies in the current layer, when the number is zero, according to the layer type of the current layer, the first policy execution path is adjusted using the third target large model, achieving efficient and accurate adjustment of the first policy execution path.
[0084] In some embodiments of the present application, the adjusting the first policy execution path based on the third target large model according to the layer type includes:
[0085] When the layer type is the layer where the root node is located, based on the third target large model, generate the third trajectory tree according to the interaction data associated with the historical policy execution path; evaluate each candidate policy in the third trajectory tree based on the third evaluation model to determine the target candidate policy in the third trajectory tree, so as to obtain the third policy execution path; use the third policy execution path to replace the first policy execution path;
[0086] When the layer type is the level where the non-root node is located, based on the third target large model, multiple new candidate strategies are generated according to the interaction data associated with the target candidate strategy in the first trajectory tree; each new candidate strategy is evaluated based on the third evaluation model to determine the target new candidate strategy, and the current target candidate strategy is replaced with the target new candidate strategy.
[0087] In this embodiment, when the layer type is the level where the root node is located, it is explained that each candidate strategy in the level where the root node in the first trajectory tree is located is not suitable, and a new trajectory tree needs to be established. Specifically, the third target large model is called from the target agent, and the third trajectory tree is generated according to the interaction data associated with the historical strategy execution path. Each candidate strategy in the third trajectory tree is a new candidate strategy or different from the candidate strategies in the level. This new candidate strategy is different from the initial strategy, thus avoiding recommending strategies that do not meet the requirements to the user again. Then, the third evaluation model is called to evaluate each candidate strategy in the third trajectory tree to determine the target candidate strategy, so as to obtain the third strategy execution path. Finally, the third strategy execution path is used to replace the first strategy execution path. Since the third target large model is different from the first target large model, it can be ensured that the third strategy execution path is different from the first strategy execution path, increasing the possibility of successfully executing the strategy.
[0088] In some embodiments of the present application, as an alternative, when the layer type is the level where the root node is located, the second trajectory tree can also be generated based on the second target large model according to the information query request and the preliminary information; each candidate strategy in the second trajectory tree is evaluated based on the second evaluation model to determine the target candidate strategy in the second trajectory tree, so as to obtain the second strategy execution path; the first strategy execution path is replaced with the second strategy execution path. Since the second target large model is different from the first target large model, it can be ensured that the second strategy execution path is different from the first strategy execution path, increasing the possibility of successfully executing the strategy.
[0089] For example, as Figure 8 shown, when the candidate strategies in the level where the node corresponding to a price decrease of x1% and a quantity increase of x2% are all executed unsuccessfully, in this case, the third trajectory tree is established by calling the third target large model and the third evaluation model, or the second trajectory tree is established by calling the second target large model and the second evaluation model. As Figure 9 shown is a schematic diagram of the second trajectory tree or the third trajectory tree in the embodiment of the present application. The structure of the second trajectory tree or the third trajectory tree and the candidate strategies corresponding to each node are different from those of the first trajectory tree.
[0090] When the layer type is the level where the non-root node is located, it indicates that all candidate nodes in the current level of the first trajectory tree do not meet the requirements, and it is necessary to re-determine the target candidate strategy in the current level. Specifically, call the third target large model from the target agent, generate multiple new candidate strategies based on the interaction data associated with the target candidate strategy, and then evaluate the generated multiple new candidate strategies by calling the third evaluation model to determine the target new candidate strategy to be recommended to the user. Finally, use the target new candidate strategy to replace the current target candidate strategy to ensure reliable strategy recommendation to the user.
[0091] For example, as Figure 10 shown, the nodes where the price drops by x1% and the quantity increases by x2% and the nodes of the prices and quantities of other products of the same type are all executed successfully, while both nodes in the level to which the nodes of the selling prices and quantities of the surrounding products belong are executed unsuccessfully. In this case, call the third target large model to generate multiple new candidate strategies based on the interaction data associated with the target candidate strategy, as Figure 11 shown, and then call the third evaluation model to evaluate the generated multiple new candidate strategies to determine the target new candidate strategy to be recommended to the user.
[0092] In some embodiments of the present application, after using the target new candidate strategy to replace the current target candidate strategy, it further includes:
[0093] Adjust the parameters of the third target large model according to the execution situation of the target new candidate strategy.
[0094] In this embodiment, after using the target new candidate strategy to replace the current target candidate strategy, determine the execution situation of the target new candidate strategy, such as whether the user executes the target new candidate strategy and whether the execution is successful, etc., and adjust the parameters of the third target large model according to the execution situation, so as to realize the automatic update of the parameters of the third target large model and improve the accuracy of the third target large model.
[0095] In some embodiments of the present application, after determining the number of executable candidate strategies in the current level, as Figure 4 shown, it further includes the following steps:
[0096] Step S1064, when the number is multiple, select a target executable candidate strategy from each of the executable candidate strategies, and use the target executable candidate strategy to replace the current target candidate strategy.
[0097] In this embodiment, when the number of executable candidate policies in the current layer is multiple, select a target executable candidate policy from other executable candidate policies in the current layer. For example, determine the executable candidate policy with the highest success probability among other executable candidate policies as the target executable candidate policy, or randomly select one from other executable candidate policies as the target executable candidate policy. Then, use the target executable candidate policy to replace the current target candidate policy, so as to accurately adjust the execution path of the first policy.
[0098] Step S1065, when the number is one, use the executable candidate policy to replace the current target candidate policy.
[0099] When the number is one, use the executable candidate policy to replace the current target candidate policy, so as to accurately adjust the execution path of the first policy.
[0100] For example, as Figure 7 shown, when the prices and quantities of other products of the same type, the nodes where the quantity is located, and the sales prices and quantities of surrounding products, such as the nodes where the computer stand is located, all fail to execute, recommend node X. Then the adjusted execution path of the first policy is the node where the price drops by x1% and the quantity increases by x2%, node X, and the nodes where the prices and quantities of other products of the same type are located.
[0101] In some embodiments of the present application, before obtaining the preliminary information of the target product, it further includes:
[0102] In response to obtaining the information query request, determine the request type of the information query request;
[0103] When the request type is a factual request, query the target knowledge graph according to the information query request to determine the first information. The target knowledge graph belongs to the target intelligent agent; recommend the first information to the user;
[0104] When the request type is a non-factual request, according to the information query request, use the fourth target model to determine the second information. The fourth target model belongs to the target intelligent agent; recommend the second information to the user;
[0105] When the second information cannot be determined using the fourth target model, execute the step of obtaining the preliminary information of the target product.
[0106] In this embodiment, a factual request is a request with an objective answer, such as a request for computer configuration information, warranty period, product material, etc. A non-factual request is a request without an objective answer, such as how to maintain a product, how to upgrade product accessories, etc. The target agent includes a target knowledge graph and a fourth target model, and the fourth target model can be based on a behavior cloning model.
[0107] In the case where the request type is a factual request, the target knowledge graph is queried according to the information query request to determine the first information and recommend it to the user, so as to provide an accurate reply to the user. In the case where the request type is a non-factual request, according to the information query request, the fourth target model is called from the target agent to determine the second information, and the second information is recommended to the user, so as to quickly reply to the user. In the case where the second information cannot be determined by using the fourth target model, the step of obtaining the preliminary information of the target product is executed, so as to ensure that corresponding strategies are reliably recommended to the user.
[0108] An embodiment of the present application also proposes a recommendation device, as Figure 12 shown, including: an obtaining module, configured to obtain preliminary information of the target product in response to obtaining an information query request of the user for the target product; a building module, configured to build a first trajectory tree according to the preliminary information and a strategy set, the strategy set includes a plurality of initial strategies, the first trajectory tree includes a plurality of levels executed in sequence, each level includes at least one candidate strategy as a node, and the candidate strategy is at least determined according to the preliminary information and the historical information corresponding to the initial strategy; a determining module, configured to determine a target candidate strategy in each level according to the evaluation results corresponding to the nodes in the first trajectory tree, and the evaluation results are determined by calling a first evaluation model to evaluate each candidate strategy; a recommendation module, configured to recommend to the user at least one first strategy execution path including the target candidate strategy.
[0109] After the recommendation device according to the embodiment of the present application obtains the information query request of the user for the target product through the obtaining module, the building module builds a first trajectory tree according to the preliminary information of the target product and the strategy set, the determining module determines the target candidate strategies in each level of the first trajectory tree according to the evaluation results corresponding to the nodes in the first trajectory tree, and then the recommendation module recommends to the user at least one first strategy execution path including the target candidate strategy. In this way, according to the information query request of the user, the strategy execution path matching the target product is automatically determined and recommended to the user, so as to realize automatically determining the product processing strategy and recommending it to the user, improving the accuracy of the processing strategy and enhancing the user experience.
[0110] In a specific application scenario, the first evaluation model belongs to the target intelligent agent, and the target intelligent agent further includes a first target large model. The establishment of the first trajectory tree based on the preliminary information and the policy set is executed based on the first target large model; or,
[0111] The target intelligent agent further includes a second target large model and a second evaluation model. The second target large model is used to generate a second trajectory tree according to the information query request and the preliminary information; the second evaluation model is used to evaluate candidate policies at multiple levels in the second trajectory tree to determine the target candidate policy, so as to obtain the second policy execution path.
[0112] In a specific application scenario, the first evaluation model belongs to the target intelligent agent, and the target intelligent agent further includes a third target large model and a third evaluation model. The third target large model is used to generate a new candidate policy according to the interaction data associated with the target candidate policy, and the new candidate policy is different from the initial policy; the third evaluation model is used to evaluate the new candidate policy; or,
[0113] In the case where the first policy execution path does not meet the requirements, the third target large model is used to generate a third trajectory tree according to the interaction data associated with the historical policy execution path. Each candidate policy in the third trajectory tree is a new candidate policy or different from the candidate policy in the level; the new candidate policy is different from the initial policy, and the historical policy execution path at least includes the first policy execution path; the third evaluation model is used to evaluate each candidate policy in the third trajectory tree.
[0114] In a specific application scenario, it further includes an adjustment module, which is used to: obtain the feedback information of the user on the first policy execution path; in the case where the feedback information meets the target condition, determine the current target candidate policy in the first policy execution path, and adjust the first policy execution path according to the execution situation of the candidate policies in the current level where the current target candidate policy is located. The target condition includes one of the following: not executing the current target candidate policy; the current target candidate policy fails to execute.
[0115] In a specific application scenario, the adjustment module is specifically used to: determine the number of executable candidate policies in the current level; in the case where the number is zero, determine the layer type of the current level, and the layer type includes the level where the root node is located or the level where the non-root node is located; based on the layer type, adjust the first policy execution path based on the third target large model.
[0116] In a specific application scenario, the adjustment module is further specifically configured to: when the layer type is the layer where the root node is located, generate the third trajectory tree based on the third target large model according to the interaction data associated with the historical policy execution path; evaluate each candidate policy in the third trajectory tree based on the third evaluation model to determine the target candidate policy in the third trajectory tree, so as to obtain the third policy execution path; use the third policy execution path to replace the first policy execution path;
[0117] When the layer type is the layer where the non-root node is located, generate a plurality of the new candidate policies based on the third target large model according to the interaction data associated with the target candidate policy in the first trajectory tree; evaluate each of the new candidate policies based on the third evaluation model to determine the target new candidate policy, and use the target new candidate policy to replace the current target candidate policy.
[0118] In a specific application scenario, the adjustment module is further configured to: adjust the parameters of the third target large model according to the execution situation of the target new candidate policy.
[0119] In a specific application scenario, the adjustment module is further configured to: when the quantity is multiple, select a target executable candidate policy from each of the executable candidate policies and use the target executable candidate policy to replace the current target candidate policy; when the quantity is one, use the executable candidate policy to replace the current target candidate policy.
[0120] In a specific application scenario, the recommendation module is further configured to: in response to obtaining the information query request, determine the request type of the information query request; when the request type is a factual request, query the target knowledge graph according to the information query request to determine the first information, and the target knowledge graph belongs to the target intelligent agent; recommend the first information to the user; when the request type is a non-factual request, determine the second information according to the information query request by using the fourth target model, and the fourth target model belongs to the target intelligent agent; recommend the second information to the user; when the second information cannot be determined by using the fourth target model, execute the step of obtaining the preliminary information of the target product.
[0121] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0122] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.
Claims
1. A recommendation method, comprising: In response to obtaining a user's information query request for a target product, obtaining preliminary information of the target product; Establishing a first trajectory tree according to the preliminary information and the strategy set, the strategy set including a plurality of initial strategies, the first trajectory tree including a plurality of levels executed sequentially, each of the levels including at least one candidate strategy as a node, the candidate strategy being determined at least according to the preliminary information and historical information corresponding to the initial strategy; Determine a target candidate strategy in each of the levels according to an evaluation result corresponding to each of the nodes in the first trajectory tree, wherein the evaluation result is determined by calling a first evaluation model to evaluate each of the candidate strategies; At least one first policy execution path including the target candidate policy is recommended to the user.
2. The recommendation method according to claim 1, wherein the first evaluation model belongs to a target agent, the target agent further comprises a first target large model, and the step of establishing the first trajectory tree according to the preliminary information and the strategy set is performed based on the first target large model; Alternatively, the target agent further includes a second target large model and a second evaluation model, The second target large model is used to generate a second trajectory tree according to the information query request and the preliminary information; The second evaluation model is used to evaluate candidate strategies at multiple levels in the second trajectory tree to determine a target candidate strategy, so as to obtain a second strategy execution path.
3. The recommendation method according to claim 1, wherein the first evaluation model belongs to a target agent, and the target agent further comprises a third target large model and a third evaluation model. The third target large model is used to generate a new candidate strategy based on the interaction data associated with the target candidate strategy, and the new candidate strategy is different from the initial strategy; The third evaluation model is used to evaluate the new candidate strategy; or, If the first strategy execution path does not meet the requirements, The third target big model is used to generate a third trajectory tree according to the interaction data associated with the historical strategy execution path, each candidate strategy in the third trajectory tree is a new candidate strategy, or is different from the candidate strategy in the hierarchy; the new candidate strategy is different from the initial strategy, and the historical strategy execution path at least includes the first strategy execution path; The third evaluation model is used to evaluate each candidate strategy in the third trajectory tree.
4. The recommendation method according to claim 3, after recommending at least one first policy execution path including the target candidate policy to the user, further comprising: Obtaining user feedback information on the execution path of the first strategy; In the case where the feedback information satisfies a target condition, determining a current target candidate strategy in the first strategy execution path, and adjusting the first strategy execution path according to the execution status of the candidate strategy in the current level where the current target candidate strategy is located, wherein the target condition includes one of the following: not executing the current target candidate strategy; The current target candidate strategy fails to execute.
5. The recommendation method according to claim 4, wherein adjusting the first strategy execution path according to the execution status of the candidate strategy in the current level where the current target candidate strategy is located comprises: determining a number of executable candidate strategies in the current level; When the number is zero, determining a layer type of the current layer, the layer type including a layer where a root node is located or a layer where a non-root node is located; According to the layer type, the first policy execution path is adjusted based on the third target big model.
6. The recommendation method according to claim 5, wherein adjusting the first strategy execution path based on the third target big model according to the layer type comprises: In the case where the layer type is the layer where the root node is located, based on the third target macro model, the third trajectory tree is generated according to the interaction data associated with the historical strategy execution path; Evaluate each candidate strategy in the third trajectory tree based on the third evaluation model, determine a target candidate strategy in the third trajectory tree, and obtain a third strategy execution path; Replacing the first policy execution path with the third policy execution path; In the case where the layer type is the layer where the non-root node is located, based on the third target large model, a plurality of new candidate strategies are generated according to the interaction data associated with the target candidate strategies in the first trajectory tree; based on the third evaluation model, each of the new candidate strategies is evaluated to determine the target new candidate strategy, and the current target candidate strategy is replaced with the target new candidate strategy.
7. The recommendation method according to claim 6, after replacing the current target candidate strategy with the new target candidate strategy, further comprising: According to the execution status of the target new candidate strategy, the parameters of the third target large model are adjusted.
8. The recommendation method according to claim 5, after determining the number of executable candidate strategies in the current level, further comprising: In the case that the number is multiple, selecting a target executable candidate strategy from each of the executable candidate strategies, and replacing the current target candidate strategy with the target executable candidate strategy; When the number is one, the current target candidate strategy is replaced by the executable candidate strategy.
9. The recommendation method according to claim 1, before obtaining the preliminary information of the target product, further comprising: In response to obtaining the information query request, determining a request type of the information query request; In the case where the request type is a factual request, querying a target knowledge graph according to the information query request to determine first information, wherein the target knowledge graph belongs to a target intelligent agent; and recommending the first information to the user; In the case where the request type is a non-factual request, determining the second information according to the information query request using a fourth target model, the fourth target model belonging to the target intelligent agent; recommending the second information to the user; In the case where the second information cannot be determined using the fourth target model, the step of obtaining preliminary information of the target product is performed.
10. A recommendation device, comprising: An obtaining module, configured to obtain preliminary information of a target product in response to a user's information query request for the target product; an establishing module, configured to establish a first trajectory tree according to the preliminary information and a strategy set, wherein the strategy set includes a plurality of initial strategies, the first trajectory tree includes a plurality of levels executed sequentially, each of the levels includes at least one candidate strategy as a node, and the candidate strategy is determined at least according to the preliminary information and historical information corresponding to the initial strategy; a determination module, configured to determine a target candidate strategy in each of the levels according to an evaluation result corresponding to each of the nodes in the first trajectory tree, wherein the evaluation result is determined by evaluating each of the candidate strategies by calling a first evaluation model; The recommendation module is used to recommend at least one first policy execution path including the target candidate policy to the user.