Event resolution strategy pushing method and device, electronic equipment and storage medium

By acquiring candidate solutions and evaluating their feasibility using the Monte Carlo algorithm, the problem of low accuracy in strategy recommendation in existing technologies is solved, achieving more accurate strategy recommendations and user solutions.

CN116501962BActive Publication Date: 2026-05-15PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2023-04-23
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, when applications push event resolution strategies, it is difficult to accurately assess the actual effect of object combinations, resulting in low strategy accuracy, large errors, and an inability to effectively solve problems for users.

Method used

By responding to event resolution requests, candidate resolution strategies are obtained, scored according to expected performance values, the feasibility of the strategies is evaluated using the Monte Carlo algorithm, and the target resolution strategy is pushed to the user when preset conditions are met.

Benefits of technology

It enables multiple evaluations of solution strategies, improves the accuracy of strategy recommendations, reduces errors, ensures the best effect of push strategies, and truly solves problems for users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116501962B_ABST
    Figure CN116501962B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for pushing an event solution strategy, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, realizes multiple evaluation of the solution strategy, greatly guarantees that the solution strategy pushed to the user is the best, improves the accuracy of strategy recommendation, reduces errors, and truly solves the problem for the user. The method comprises the following steps: in response to an event solution request, determining a to-be-solved event, and acquiring a plurality of candidate solution strategies associated with the to-be-solved event; scoring each candidate solution strategy according to an expected effect value corresponding to each candidate solution strategy in the plurality of candidate solution strategies, to obtain a strategy score value corresponding to each candidate solution strategy; extracting a target solution strategy with the maximum strategy score value from the plurality of candidate solution strategies, and evaluating the feasibility by using a Monte Carlo algorithm to obtain a feasibility conclusion; and when the feasibility conclusion meets a preset condition, determining a user who initiates the event solution request and pushing the target solution strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for pushing out an event resolution strategy. Background Technology

[0002] With the continuous development of the internet and artificial intelligence technologies, applications providing users with a wide variety of online services are emerging. For example, medical applications allow users to consult doctors, schedule follow-up appointments, purchase medications, and more online, eliminating the hassle of in-person visits. When a user requests a service on the application, the application treats the request as an event and selects multiple solutions for that event, pushing these solutions to the user for selection. For instance, during a follow-up appointment, the application intelligently recommends several available medications based on the user's current condition, allowing the user to choose and purchase the necessary medication.

[0003] In related technologies, when applications push solutions to events, they simultaneously present multiple solutions to the user. The applicant recognizes that these strategies have varying effectiveness, and some strategies actually involve combinations of medications; for example, for a specific illness, the application might push a strategy involving a combination of drugs. However, when evaluating strategies, applications typically rely solely on historical user feedback or which strategies were most frequently adopted. This results in a somewhat one-sided evaluation, making it difficult to accurately assess the actual effects of combinations of medications in strategies. Consequently, the accuracy of the strategies pushed to users is low, with significant errors, and they fail to truly solve the user's problem. Summary of the Invention

[0004] In view of this, this application provides a method, apparatus, electronic device and storage medium for pushing event resolution strategies. The main purpose is to solve the problem that it is currently difficult to truly evaluate the actual effect of the combination of objects, the strategy pushed to the user has low accuracy and large error, and cannot truly solve the problem for the user.

[0005] According to a first aspect of this application, a method for pushing out an event resolution strategy is provided, the method comprising:

[0006] In response to an event resolution request, determine the event to be resolved indicated by the event resolution request, and obtain multiple candidate resolution strategies associated with the event to be resolved;

[0007] Based on the expected effect value of each candidate solution strategy among the multiple candidate solution strategies, each candidate solution strategy is scored to obtain the strategy score value corresponding to each candidate solution strategy.

[0008] The candidate solution strategy with the largest strategy score is selected as the target solution strategy from the multiple candidate solution strategies, and the feasibility of the target solution strategy is evaluated using the Monte Carlo algorithm to obtain a feasibility conclusion for the target solution strategy.

[0009] When the feasibility conclusion meets the preset conditions, the user who initiated the event resolution request is identified, and the target resolution strategy is pushed to the user's terminal.

[0010] Optionally, obtaining multiple candidate resolution strategies associated with the event to be resolved includes:

[0011] Obtain the event identifier of the event to be resolved, query multiple optional result values ​​associated with the event identifier, and use each of the multiple optional result values ​​as a candidate resolution strategy to obtain the multiple candidate resolution strategies; or,

[0012] A preset number of combinations is determined, and the multiple optional result values ​​are arranged and combined according to the preset number of combinations to obtain multiple result value arrays. Each result value array in the multiple result value arrays is used as a candidate solution strategy to obtain the multiple candidate solution strategies. The number of optional result values ​​included in each result value array is equal to the preset number of combinations.

[0013] Optionally, the step of scoring each candidate solution strategy based on the expected effect value corresponding to each candidate solution strategy among the plurality of candidate solution strategies, to obtain a strategy score value corresponding to each candidate solution strategy, includes:

[0014] For each of the plurality of candidate solution strategies, a first optional result value is read from the candidate solution strategy;

[0015] Determine the method for obtaining the first optional result value, and query the probability of obtaining the first optional result value based on the method of obtaining it;

[0016] Obtain the expected effect value associated with the first optional result value, calculate the product of the expected effect value and the acquisition probability, and use the product as the strategy score corresponding to the candidate solution strategy.

[0017] Optionally, the step of using the Monte Carlo algorithm to evaluate the feasibility of the target solution strategy and obtaining a feasibility conclusion for the target solution strategy includes:

[0018] The second optional result value read in the target solution strategy is used as the target acquisition method;

[0019] The target acquisition method is simulated, and during the simulation, the actions that occur are collected, and the preset action scores associated with the actions that occur are determined. The simulation result value is obtained after the simulation is completed.

[0020] Obtain the expected reward function constructed using the Monte Carlo algorithm, input the action value of the action and the preset action score into the expected reward function, and use the output value of the expected reward function as the expected reward value for this simulation process;

[0021] The simulation result value is correlated with the expected return value and used as the simulation round result of this simulation process, and the simulation round result is recorded;

[0022] The target acquisition method is re-simulated and the simulation round results are regenerated until the number of simulations reaches a first preset number. The first preset number of simulation round results are obtained, and the first preset number of simulation round results are used as the feasibility conclusion of the target solution strategy.

[0023] Optionally, after selecting the candidate solution strategy with the largest strategy score from the plurality of candidate solution strategies as the target solution strategy, and evaluating the feasibility of the target solution strategy using the Monte Carlo algorithm to obtain a feasibility conclusion for the target solution strategy, the method further includes:

[0024] The preset conditions are obtained, and the preset conditions include standard result values, standard return values, and quantity thresholds;

[0025] The feasibility conclusion reads a first preset number of simulation round results, and counts the number of results for the target simulation round. The target simulation round results are the simulation round results in the first preset number of simulation round results where the simulation result value is equal to the standard result value and the expected return value is greater than or equal to the standard return value.

[0026] The number of results is compared with the number threshold;

[0027] Accordingly, if the number of results is greater than or equal to the number threshold, then the feasibility conclusion is determined to meet the preset condition;

[0028] If the number of results is less than the number threshold, it is determined that the feasibility conclusion does not meet the preset condition. Among the multiple candidate solution strategies, the remaining candidate solution strategies other than the target solution strategy are determined. Among the remaining candidate solution strategies, the candidate solution strategy with the largest strategy score is extracted as the designated solution strategy. The feasibility of the designated solution strategy is evaluated using the Monte Carlo algorithm to obtain the feasibility conclusion of the designated solution strategy, and it is determined whether the feasibility conclusion of the designated solution strategy meets the preset condition.

[0029] Optionally, the method further includes:

[0030] If there are no associated candidate solutions for the event to be resolved, then multiple state variables are constructed;

[0031] The Monte Carlo algorithm is used to simulate the execution of the event to be resolved according to the possible values ​​of each of the plurality of state variables;

[0032] Obtain the target possible values ​​corresponding to each state variable in this simulation, and label each state variable using the target possible values ​​corresponding to each state variable;

[0033] Obtain the simulation reward value after the simulation ends, associate the labeled multiple state variables with the simulation reward value to obtain a simulation data set, and record the simulation data set;

[0034] The Monte Carlo algorithm is re-adopted, and the event to be resolved is simulated and executed according to the possible values ​​of each of the multiple state variables. The new possible values ​​of the target corresponding to each state variable are obtained and labeled. The new simulation reward value obtained in this simulation is associated with the relabeled multiple state variables to generate a new simulation data set. The simulation is repeated until the number of rounds reaches the second preset number, and the second preset number of simulation data sets are obtained.

[0035] Extract the target simulation data group with the largest simulated return value from the second preset number of simulation data groups, combine the multiple possible values ​​corresponding to the multiple state variables included in the target simulation data group to obtain a value group, and use the data group as the optimal solution strategy for the event to be solved.

[0036] The optimal solution strategy is pushed to the user's terminal.

[0037] Optionally, the step of employing the Monte Carlo algorithm to simulate the event to be resolved according to the possible values ​​of each of the plurality of state variables includes:

[0038] Based on the possible values ​​corresponding to each state variable, a possible value is randomly selected for each state variable, and the selected possible value is used to assign a value to the corresponding state variable to obtain the multiple state variables after assignment.

[0039] According to the order of change of the multiple state variables, determine the first state variable and the next state variable of the first state variable among the multiple state variables;

[0040] Simulate the process of transitioning from the first state variable to the next state variable, acquire simulated action values ​​during the transition, and determine the predicted transition probability of the transition from the first state variable to the next state variable;

[0041] Obtain the action value function constructed using the Monte Carlo algorithm, input the simulated action value and the predicted transition probability into the action value function, and use the output value of the action value function as the simulated operation value;

[0042] Continue to determine the next specified state variable of the next state variable among the multiple state variables, and calculate the simulation operation value of the next state variable and the specified state variable, until the multiple state variables are traversed, and end the current simulation, and obtain multiple simulation operation values;

[0043] The simulation operation value with the largest value among the multiple simulation operation values ​​is taken as the simulation reward value after the end of this simulation.

[0044] According to a second aspect of this application, a device for pushing out an event resolution strategy is provided, the device comprising:

[0045] The acquisition module is used to respond to an event resolution request, determine the event to be resolved indicated by the event resolution request, and acquire multiple candidate resolution strategies associated with the event to be resolved;

[0046] The scoring module is used to score each candidate solution strategy based on the expected effect value corresponding to each candidate solution strategy among the multiple candidate solution strategies, and obtain the strategy score value corresponding to each candidate solution strategy.

[0047] The evaluation module is used to extract the candidate solution strategy with the largest strategy score from the multiple candidate solution strategies as the target solution strategy, and to evaluate the feasibility of the target solution strategy using the Monte Carlo algorithm to obtain a feasibility conclusion of the target solution strategy.

[0048] The push module is used to determine the user who initiated the event resolution request when the feasibility conclusion meets the preset conditions, and to push the target resolution strategy to the terminal held by the user.

[0049] Optionally, the acquisition module is configured to acquire the event identifier of the event to be resolved, query multiple optional result values ​​associated with the event identifier, and use each of the multiple optional result values ​​as a candidate resolution strategy to obtain the multiple candidate resolution strategies; or, determine a preset number of combinations, arrange and combine the multiple optional result values ​​according to the preset number of combinations to obtain multiple result value arrays, and use each of the multiple result value arrays as a candidate resolution strategy to obtain the multiple candidate resolution strategies, wherein the number of optional result values ​​included in each result value array is equal to the preset number of combinations.

[0050] Optionally, the scoring module is configured to, for each of the plurality of candidate solution strategies, read a first optional result value from the candidate solution strategy; determine the acquisition method of the first optional result value, query the acquisition probability of the first optional result value based on the acquisition method; acquire the expected effect value associated with the first optional result value, calculate the product of the expected effect value and the acquisition probability, and use the product as the strategy score value corresponding to the candidate solution strategy.

[0051] Optionally, the evaluation module is configured to: read a second optional result value from the target solution strategy; use the method of obtaining the second optional result value as the target acquisition method; simulate the execution of the target acquisition method, collect the actions that occur during the simulation, and determine the preset action scores associated with the actions that occur; obtain the simulation result value obtained after the simulation is completed; obtain the expected reward function constructed using the Monte Carlo algorithm; input the action values ​​of the actions that occur and the preset action scores into the expected reward function; use the output value of the expected reward function as the expected reward value for this simulation process; associate the simulation result value with the expected reward value and use it as the simulation round result for this simulation process; record the simulation round result; re-simulate the execution of the target acquisition method and regenerate the simulation round result for the simulation process until the number of simulations reaches a first preset number, obtain the first preset number of simulation round results, and use the first preset number of simulation round results as the feasibility conclusion of the target solution strategy.

[0052] Optionally, the evaluation module is further configured to: acquire the preset conditions, extract the preset conditions including a standard result value, a standard return value, and a quantity threshold; read a first preset number of simulation round results from the feasibility conclusion, count the number of results for the target simulation round, wherein the target simulation round result is the simulation round result in the first preset number of simulation rounds where the simulation result value is equal to the standard result value and the expected return value is greater than or equal to the standard return value; compare the number of results with the quantity threshold; accordingly, if the number of results is greater than or equal to the quantity threshold, then determine that the feasibility conclusion satisfies the preset conditions; wherein, if the number of results is less than the quantity threshold, then determine that the feasibility conclusion does not satisfy the preset conditions; determine the remaining candidate solution strategies other than the target solution strategy from the multiple candidate solution strategies; extract the candidate solution strategy with the largest strategy score from the remaining candidate solution strategies as the designated solution strategy; and evaluate the feasibility of the designated solution strategy using the Monte Carlo algorithm to obtain the feasibility conclusion of the designated solution strategy, and determine whether the feasibility conclusion of the designated solution strategy satisfies the preset conditions.

[0053] Optionally, the device further includes:

[0054] A construction module is used to construct multiple state variables if there are no associated candidate resolution strategies for the event to be resolved.

[0055] The simulation module is used to simulate the execution of the event to be solved by employing the Monte Carlo algorithm according to the possible values ​​of each of the plurality of state variables;

[0056] The annotation module is used to obtain the target possible values ​​corresponding to each state variable in this simulation, and to annotate each state variable using the target possible values ​​corresponding to each state variable.

[0057] The association module is used to obtain the simulation reward value obtained after the simulation ends, associate the labeled multiple state variables with the simulation reward value to obtain a simulation data group, and record the simulation data group;

[0058] The simulation module is also used to re-emulate the Monte Carlo algorithm, simulate the execution of the event to be resolved according to the possible values ​​of each of the multiple state variables, and re-obtain and label the new possible values ​​of each state variable, associate the new simulation reward value obtained in this simulation with the re-labeled multiple state variables, generate a new simulation data set, until the number of simulation rounds reaches the second preset number, and obtain the second preset number of simulation data sets;

[0059] The extraction module is used to extract the target simulation data group with the largest simulation return value from the second preset number of simulation data groups, combine multiple possible values ​​corresponding to multiple state variables included in the target simulation data group to obtain a value group, and use the data group as the optimal solution strategy for the event to be solved.

[0060] The push module is also used to push the optimal solution strategy to the user's terminal.

[0061] Optionally, the simulation module is configured to: randomly select a possible value for each state variable based on the possible values ​​corresponding to each state variable; assign the selected possible value to the corresponding state variable to obtain the assigned multiple state variables; determine the first state variable and the next state variable of the first state variable from the multiple state variables according to the change order corresponding to the multiple state variables; simulate the transition from the first state variable to the next state variable, obtain simulated action values ​​during the transition, and determine the predicted transition probability from the first state variable to the next state variable; obtain an action value function constructed using the Monte Carlo algorithm, input the simulated action value and the predicted transition probability into the action value function, and use the output value of the action value function as the simulated operation value; continue to determine the next specified state variable from the multiple state variables, and calculate the simulated operation value of the next state variable and the specified state variable, until the multiple state variables are traversed, ending the current simulation and obtaining multiple simulated operation values; and use the simulation operation value with the largest value among the multiple simulated operation values ​​as the simulation reward value after the end of the current simulation.

[0062] According to a third aspect of this application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in any of the first aspects above.

[0063] According to a fourth aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the steps of the method described in any one of the first aspects above.

[0064] By employing the above technical solutions, this application provides a method, apparatus, electronic device, and storage medium for pushing event resolution strategies. In response to an event resolution request, this application determines the event to be resolved indicated by the event resolution request, obtains multiple candidate resolution strategies associated with the event, scores each candidate resolution strategy based on the expected effect value corresponding to each candidate resolution strategy, obtains a strategy score value for each candidate resolution strategy, extracts the candidate resolution strategy with the highest strategy score value from the multiple candidate resolution strategies as the target resolution strategy, and uses a Monte Carlo algorithm to evaluate the feasibility of the target resolution strategy, obtaining a feasibility conclusion for the target resolution strategy. When the feasibility conclusion meets preset conditions, it determines the user who initiated the event resolution request and pushes the target resolution strategy to the user's terminal. This achieves multiple evaluations of the resolution strategy, ensuring to a large extent that the resolution strategy pushed to the user has the best effect, improving the accuracy of strategy recommendations, reducing errors, and truly solving problems for users.

[0065] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0066] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0067] Figure 1 This illustration shows a flowchart of a push method for an event resolution strategy provided in an embodiment of this application;

[0068] Figure 2A This illustration shows a flowchart of a push method for another event resolution strategy provided in an embodiment of this application;

[0069] Figure 2B A schematic diagram of a push method for an event resolution strategy provided in an embodiment of this application is shown;

[0070] Figure 2C A schematic diagram of a push method for an event resolution strategy provided in an embodiment of this application is shown;

[0071] Figure 3 This illustration shows a schematic diagram of the structure of a push device for an event resolution strategy provided in an embodiment of this application;

[0072] Figure 4A schematic diagram of the device structure of a computer device provided in an embodiment of this application is shown. Detailed Implementation

[0073] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.

[0074] This application provides a method for pushing event resolution strategies, such as... Figure 1 As shown, the method includes:

[0075] 101. In response to an event resolution request, determine the event to be resolved as indicated by the event resolution request, and obtain multiple candidate resolution strategies associated with the event to be resolved.

[0076] In real-world scenarios, some situations may have multiple solutions. For example, patients have many options when purchasing medication for their first or follow-up visit. However, the applicant recognizes that some available medications, when used in combination, have a greater effect, while others, when combined, are not as effective. Finding the strategy to maximize the medication's effectiveness has become a pressing issue.

[0077] Therefore, this application proposes an event resolution strategy. In response to an event resolution request, multiple candidate resolution strategies are first scored, and an attempt is made to derive a target resolution strategy that is guaranteed to be optimal in most cases. Then, the Monte Carlo algorithm is used to evaluate the target resolution strategy. When the feasibility of the target resolution strategy meets preset conditions, it is output as the resolution strategy recommended to the user. This achieves multiple evaluations of the resolution strategy, ensuring to a large extent that the resolution strategy pushed to the user has the best effect, improving the accuracy of strategy recommendation, reducing errors, and truly solving problems for the user.

[0078] The technical solution of this application embodiment can be applied to a platform providing medical services. The platform provides a front-end application for users, where they can register for appointments, schedule physical examinations, consult doctors online, and schedule follow-up visits. When a user has a problem that needs to be resolved, they can report the problem as a pending event to the platform. The platform can then detect the event resolution request, identify the pending event indicated by the request, and obtain multiple candidate resolution strategies associated with the pending event for initial evaluation. Furthermore, the platform relies on the computing power of servers to provide services to users. These servers can be independent servers or servers providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, etc., to provide users with services such as physical examinations, online medical consultations, and medical record viewing. On the platform, users can browse medical data, electronic medical records, health check packages, and more. Specifically, medical data can include personal health records, prescription information, examination reports, and other data; electronic medical records are electronic healthcare records, which can include a series of electronic personal health records with archival value, such as medical records, electrocardiograms, and medical images; and health check packages can be the health check items or combinations of health check items that can be purchased by the browser.

[0079] Furthermore, the multiple candidate solutions associated with the unresolved event can be several strategies currently available to the user. For example, assuming the unresolved event is the user's follow-up medical examination result, the candidate solutions could be multiple medications or combinations of medications that can be used for that result. In practical applications, the platform may also offer users dice games, number-matching games, etc., for relaxation. Therefore, the unresolved event could also be a user requesting the platform to provide a winning strategy while playing a dice game. In this case, the platform will determine multiple candidate solutions based on the numbers on each face of the dice, and then select the solution with the highest benefit from these options to recommend to the user.

[0080] 102. Based on the expected effect value of each candidate solution strategy among multiple candidate solution strategies, score each candidate solution strategy to obtain the strategy score value corresponding to each candidate solution strategy.

[0081] In the embodiments of this application, each candidate solution strategy corresponds to an expected effect value. The expected effect value actually indicates how much effect the corresponding candidate solution strategy can bring after application. For example, if the candidate solution strategy is a drug combination, the corresponding expected effect value indicates the effect value that the user can bring after using the drug combination; if the candidate solution strategy is a dice roll combination, the corresponding expected effect value indicates how many points can be obtained by choosing the roll combination.

[0082] In order to make a comprehensive evaluation of each candidate solution strategy, this embodiment of the application will score each candidate solution strategy according to the expected effect value of each candidate solution strategy among multiple candidate solution strategies, and obtain the strategy score value corresponding to each candidate solution strategy.

[0083] 103. Extract the candidate solution strategy with the largest strategy score from multiple candidate solution strategies as the target solution strategy, and use the Monte Carlo algorithm to evaluate the feasibility of the target solution strategy to obtain a feasibility conclusion.

[0084] In this embodiment, after obtaining the strategy score for each candidate solution strategy, the platform extracts the candidate solution strategy with the highest strategy score from among multiple candidate solutions as the target solution strategy. Considering that there may be errors in this score calculation, the platform will continue to use the Monte Carlo algorithm to evaluate the target solution strategy, determine whether the target solution strategy is truly feasible, obtain a feasibility conclusion for the target solution strategy, and subsequently determine whether to push the target solution strategy to the user based on the feasibility conclusion.

[0085] 104. When the feasibility conclusion meets the preset conditions, identify the user who initiated the event resolution request and push the target resolution strategy to the user's terminal.

[0086] In this embodiment, to evaluate the feasibility of a solution strategy, the platform is configured with preset conditions. These preset conditions are used to assess whether the result value and reward value of the solution strategy are sufficiently large. When the feasibility conclusion meets the preset conditions, it indicates that the target solution strategy is good enough and relatively ideal, and can be recommended to the user. Therefore, the platform identifies the user who initiated the event resolution request and pushes the target solution strategy to the user's terminal.

[0087] The method provided in this application, in response to an event resolution request, determines the event to be resolved indicated by the event resolution request, obtains multiple candidate resolution strategies associated with the event to be resolved, scores each candidate resolution strategy according to the expected effect value corresponding to each candidate resolution strategy, obtains a strategy score value corresponding to each candidate resolution strategy, extracts the candidate resolution strategy with the largest strategy score value from the multiple candidate resolution strategies as the target resolution strategy, and uses a Monte Carlo algorithm to evaluate the feasibility of the target resolution strategy to obtain a feasibility conclusion of the target resolution strategy. When the feasibility conclusion meets preset conditions, it determines the user who initiated the event resolution request and pushes the target resolution strategy to the user's terminal, realizing multiple evaluation of the resolution strategy, ensuring to a large extent that the resolution strategy pushed to the user has the best effect, improving the accuracy of strategy recommendation, reducing errors, and truly solving problems for users.

[0088] Furthermore, as a refinement and extension of the specific implementation methods of the above embodiments, in order to fully illustrate the specific implementation process of this embodiment, this application provides another method for pushing event resolution strategies, such as... Figure 2A As shown, the method includes:

[0089] 201. In response to an event resolution request, determine the event to be resolved as indicated by the event resolution request, and obtain multiple candidate resolution strategies associated with the event to be resolved.

[0090] The technical solution of this application embodiment can be applied to a platform providing medical services. The platform provides a front-end application to users, where users can register, schedule physical examinations, consult doctors online, and have follow-up visits. When a user has a problem that needs to be solved, the user can report the problem as an unresolved event to the platform. In this way, the platform can detect the event resolution request, determine the unresolved event indicated by the event resolution request, and obtain multiple candidate resolution strategies associated with the unresolved event, so as to conduct an initial evaluation of these candidate resolution strategies.

[0091] Specifically, when obtaining multiple candidate resolution strategies associated with an event to be resolved, the following two methods can be used: One method is to obtain the event identifier of the event to be resolved, query the multiple optional result values ​​associated with the event identifier, and use each of the multiple optional result values ​​as a candidate resolution strategy, thus obtaining multiple candidate resolution strategies. The other method is to determine a preset number of combinations, arrange and combine the multiple optional result values ​​according to the preset number of combinations, obtain multiple result value arrays, and use each result value array as a candidate resolution strategy, thus obtaining multiple candidate resolution strategies. Here, the number of optional result values ​​included in each result value array is equal to the preset number of combinations.

[0092] Because this application involves many calculations and a large number of numerical values, for ease of description, the following explanation will use a dice game provided by the platform as an example. Specifically, the dice game described is Cubilete (a traditional dice game), where the user initially rolls 5 dice, up to 3 times. After each roll, the user keeps the dice they want and rolls the remaining dice; a 1 can be considered any other number. After the game ends, rolling 5 1s earns 10 points; rolling 5 6s (excluding 1s) earns 5 points; rolling 5 6s (including 1s) earns 2 points; rolling 5 unique numbers earns 1 point; and all other cases earn 0 points. Therefore, in this embodiment, the multiple candidate solutions obtained are combinations of the numbers 1-6 on the six sides of the dice. Accordingly, in practical applications, since the effects of each drug vary and are derived from historical user feedback, the platform can pre-assign drugs based on their effects, or score them according to user evaluations. This assigns each drug a numerical value, similar to a dice roll, facilitating subsequent calculations of various parameters. For example, in treating migraines, many users report that lamotrigine is more effective than montelukast sodium. Lamotrigine could be assigned a number 6, and montelukast sodium a number 5, making it easier to calculate which is more suitable for the user's current condition.

[0093] 202. Based on the expected effect value of each candidate solution strategy among multiple candidate solution strategies, score each candidate solution strategy to obtain the strategy score value corresponding to each candidate solution strategy.

[0094] In this embodiment, after determining multiple candidate solutions, the platform scores each candidate solution based on its expected performance value, thus obtaining a strategy score for each candidate solution. The specific scoring process is as follows:

[0095] Taking a dice game as an example, typically, a user will keep a 1 in their last roll because it can be considered any value; rolling five 1s returns the highest score. However, in practice, some specific situations need to be considered separately. Since there is no penalty in this game, and the opponent's performance is disregarded, the user's goal is to obtain the highest possible score. Once the dice are kept, they cannot be rolled again. Therefore, the actions in the next two rounds are determined by the actions in the first round. In the general simulation, the decision made in the first round is called the "strategy." Therefore, to simplify the above process, this application divides the results of the first round into general cases and special cases. The special cases do not include 1s or five 1s, while the general cases consider all other cases. For special cases, since the probabilities of getting [1,1,1,1,1,] and [6,6,6,6,6,] are the least likely, they are only considered when the number of 1s or 6s is greater than 3. [2,3,3,4,5] is an example of the general case because the probability of getting close to [1,1,1,1,1,] and [6,6,6,6,6,] in the last roll is negligible. The first roll result [1,1,1,1,1,] is a typical example of the special case because the probability of producing the highest result [1,1,1,1,1,] is as high as 11 / 36. In general, rolling a 5 is considered the simplest way to score. Based on the above assumptions, this application will retain all occurrences of 1s. To avoid getting 0 points, only one number can be chosen from the 5 numbers. Thus, the dice can be further divided into two categories: [6] and the set [2,3,4,5]. In summary, the general strategy has been simplified to choosing between keeping [6] or keeping the set [2,3,4,5].

[0096] In general, it is necessary to compare the expected policy score values ​​to determine whether to maintain the pattern of [6] or the set [2,3,4,5]. The expected policy score value can be defined using the following formula 1:

[0097] Formula 1: E=∑a∈AR(a)×P(a)

[0098] Where x is the candidate solution strategy, E is the score of the corresponding strategy, A is the set of all possible final results under strategy x, and R is the expected effect value. In practical applications, all results with a score of 0 can be deleted. Thus, under normal circumstances, the above formula 1 can be redefined as the following formula 2:

[0099] Formula 2: E = R(x wins) * P(x wins)

[0100] Where x is the strategy of leaving a 6 in the first roll and leaving the other roll starting from the first roll. For the probability of winning P(x wins), a new variable n is introduced, representing the number of dice remaining after the roll. For example, suppose the result of the first roll is [1, 3, 3, 4, 6]. If 1 and 6 are kept, then n = 3. When n is known, the probability of approaching the last roll can be calculated. Thus, a variable is set to evaluate the expected reward value of different strategies. The strategy score under strategy x can be defined using the following formula 3:

[0101] Formula 3: E = R(x wins) * P(score a | choose strategy x, N = n)

[0102] Thus, as can be seen from the above, when calculating the strategy score for each candidate solution strategy among multiple candidate solution strategies, it is necessary to read the first optional result value from the candidate solution strategy, which is the possible number to be rolled. Then, determine the acquisition method of the first optional result value, query the acquisition probability of the first optional result value based on the acquisition method, which is the aforementioned P, and obtain the expected effect value associated with the first optional result value, which is the aforementioned R. Calculate the product of the expected effect value and the acquisition probability, and use the product as the strategy score corresponding to the candidate solution strategy.

[0103] 203. Extract the candidate solution strategy with the largest strategy score from multiple candidate solution strategies as the target solution strategy, and use the Monte Carlo algorithm to evaluate the feasibility of the target solution strategy to obtain a feasibility conclusion.

[0104] In this embodiment, the platform extracts the candidate solution strategy with the highest strategy score from multiple candidate solution strategies as the target solution strategy, and then uses the Monte Carlo algorithm to evaluate the feasibility of the target solution strategy to obtain a feasibility conclusion. Specifically, to use the Monte Carlo algorithm for strategy evaluation, this application pre-constructs an expected reward function using the Monte Carlo algorithm, as detailed below:

[0105] For the objective-solving strategy, this application considers using the Monte Carlo method to find the state-value function, specifically V, which needs to be found from empirical fragments under policy π. π Value functions, S1, A1, R2, ..., S k~π , where S i Represents state i, A in the experience fragment. i Representative action, R i The expected effect value represents the state. Thus, by organizing the above content, the final score formula is as follows: Formula 4:

[0106] Formula 4: Gt =R t+1 +γR t+2 +…+γT-1R T

[0107] Accordingly, let V π (s) represents the state value under the policy, and π represents the expected return, which is the expected cumulative future return starting from the initial state. Therefore, V is calculated. π When (s), use the following formula 5:

[0108] Formula 5: V π (s)=E π [G t |S t =s]

[0109] It should be noted that a simple way to estimate empirical returns is to average the returns observed afterward, starting from the initial state and visiting all states. Suppose we now need to estimate V... π (s) Given that each occurrence of state s from the current π is considered a visit to s in one round. s may be visited multiple times in the same round, such as... Figure 2B As shown, this is referred to as the first time visit round. There are two methods for estimating V during the first visit: the first visit MC method and the subsequent visit MC method. π (s). The average return s of the first visit MC method is the return s of the first visit during all events in this strategy, while the average return s of the MC method for each visit is the return s of all visits. Because this application focuses on the first visit MC method, the evaluation concept of the first visit MC strategy is detailed below, and the core idea and structure of the algorithm are as follows:

[0110] In this algorithm, the input is a candidate solution strategy π to be evaluated. First, initialization is performed. During initialization, V(s)∈R, and for any s∈S, Returns(s)← is an empty set. The algorithm iterates through all s∈S, generating one round under strategy π: S0, A0, R1, S1, A1, R2, ..., S T-1 A T-1 R T G←0. Then, for each step in a loop, t = T-1, T-2, ..., 0: G←G+R t+1 Unless S t Appearing in S0, S1, ... S t-1 Add G to Returns(S) t ), V(S t ) ← Average (Returns(S) t )).

[0111] Thus, the method converges to V. π(s) becomes infinite as the number of visits (first visit) s. This is because, in this case, each reward is an independent and identically distributed estimate V. π (s) have finite variance. According to the law of large numbers, the sequence of these means converges to their expected values. Each mean is an unbiased estimate with a standard deviation of 1 / √n for its error term, where n is the total average return over the entire empirical period. Furthermore, the above algorithm can be updated using incremental mean calculation. The means μ1 and μ2 of the sequence x1, x2, ... can be calculated incrementally, as shown in Formula 6 below:

[0112] Formula 6:

[0113] Thus, after updating the Monte Carlo algorithm using the incremental mean calculation method, V π (s) After each round, incrementally update S1, A1, R2, ..., ST, with a return Gt for each state St, expressed as: N(St)←N(St)+1, V(St)←(V(St)+1 / N(St))*(Gt-V(St)). In non-stationary problems, tracking the running average (i.e., forgetting previous rounds) can be useful, V(St)←V(St)+α(Gt-V(St)).

[0114] Thus, using the above process, this application finds the expected reward function constructed using the Monte Carlo algorithm, namely Formula 5. Next, the second optional result value read from the target solution strategy is used as the target acquisition method, and the target acquisition method is simulated. During the simulation, actions are collected, and preset action scores associated with those actions are determined. The simulation result value obtained after the simulation is completed is obtained, and the expected reward function constructed using the Monte Carlo algorithm is obtained. The action values ​​of the actions and the preset action scores are input into the expected reward function, and the output value of the expected reward function is used as the expected reward value for this simulation process. The simulation result value and the expected reward value are correlated and used as the simulation round result for this simulation process, and the simulation round result is recorded. To ensure the accuracy of the results, this application will re-simulate the target acquisition method and regenerate the simulation round result until the number of simulations reaches a first preset number, obtaining the first preset number of simulation round results. The first preset number of simulation round results are used as the feasibility conclusion of the target solution strategy.

[0115] 204. Evaluate the feasibility of the target solution strategy based on preset conditions. If the feasibility of the solution meets the preset conditions, proceed to step 205. If the feasibility of the solution does not meet the preset conditions, proceed to step 206.

[0116] In this embodiment of the application, to evaluate the feasibility of a solution strategy, the platform is equipped with preset conditions. These preset conditions are used to evaluate whether the result value and reward value of the solution strategy are sufficiently large. The principle of the feasibility assessment conclusion is described below:

[0117] In practical applications, the action is a decision made between two rolls. Regarding the state during the experience, this application defines seven state variables (λ, x1, x2, x3, x4, x5, x6) to describe the current state. During the simulation, this application sets the expected reward value between the first two dice rolls to 0, and the number of points in the user's last sequence is the return value for this round. Here, this application uses the first-visit MC strategy evaluation to test whether the candidate solution strategy is applicable to the actual situation. Specifically, the evaluation process can simulate 500,000 times, and the value function changes in each simulation need to asymptotically converge to the truth function according to the law of large numbers. This method will provide a series of values ​​for different states at different times throughout the entire experience under this strategy. Thus, once the result table is obtained, it can be determined whether the candidate solution strategy is consistent with the values ​​in the table.

[0118] For example, suppose there are candidate solution strategies π: compare the number of n6 and n+1, and then choose different methods. The results table is summarized, revealing n, n6, and the corresponding state values. Next, a value-state function table (Table 1 below) is created to more intuitively explain this result. Based on the strategy assumption π, some results from Volume 1 (λ = 1) are given:

[0119]

[0120]

[0121] Table 1

[0122] As shown in Table 1 above, there are no zero values, indicating that the candidate solution strategy has some feasibility. The basic idea behind the winning strategy is to obtain the score with the highest probability. If this candidate solution strategy is good enough, the ideal value for state 1 should be at least 1. However, most values ​​are less than 1. In other words, although this candidate solution strategy does increase the probability of scoring, the optimization effect is very small when winning.

[0123] For further examples, please refer to the obtained one-valued state function table, namely Table 2 below:

[0124]

[0125] Table 2

[0126] As can be seen from Table 2 above, these values ​​are much larger than those in Table 1, with some exceeding the minimum level score of 1. Furthermore, the state value shows an increasing trend with the increase in the number of points, consistent with the initial analysis of point 1 in this application.

[0127] In summary, based on a series of analyses, this candidate solution strategy can be considered a relatively optimal strategy in a general sense.

[0128] The aforementioned evaluation process can be simplified to obtaining preset conditions, which include standard result values, standard return values, and quantity thresholds. Subsequently, the platform reads a first preset number of simulated round results from the feasibility conclusion, counts the number of results for the target simulated round, and defines the target simulated round as the simulated round whose simulated result value equals the standard result value and whose expected return value is greater than or equal to the standard return value among the first preset number of simulated rounds. The number of results is then compared to the quantity threshold. Accordingly, if the number of results is greater than or equal to the quantity threshold, the feasibility conclusion is determined to meet the preset conditions; otherwise, if the number of results is less than the quantity threshold, the feasibility conclusion is determined not to meet the preset conditions.

[0129] 205. When the feasibility conclusion meets the preset conditions, identify the user who initiated the event resolution request and push the target resolution strategy to the user's terminal.

[0130] In this embodiment, when the feasibility conclusion meets preset conditions, it indicates that the target solution strategy is good enough and relatively ideal, and can be recommended to the user. Therefore, the platform identifies the user who initiated the event resolution request and pushes the target solution strategy to the user's terminal. Specifically, the target solution strategy can be described in detail and sent to the user's email address, account, etc.

[0131] 206. When the feasibility conclusion does not meet the preset conditions, determine the remaining candidate solutions other than the target solution strategy among multiple candidate solutions. Extract the candidate solution strategy with the largest strategy score from the remaining candidate solutions strategy as the specified solution strategy. Use the Monte Carlo algorithm to evaluate the feasibility of the specified solution strategy, obtain the feasibility conclusion of the specified solution strategy, and determine whether the feasibility conclusion of the specified solution strategy meets the preset conditions.

[0132] In this embodiment of the application, when the feasibility conclusion does not meet the preset conditions, it indicates that the target solution strategy is not good enough and other candidate solutions need to be selected and evaluated again. Therefore, the platform will determine the remaining candidate solutions other than the target solution strategy from multiple candidate solutions, extract the candidate solution strategy with the largest strategy score from the remaining candidate solutions as the designated solution strategy, and use the Monte Carlo algorithm to evaluate the feasibility of the designated solution strategy to obtain the feasibility conclusion of the designated solution strategy, and determine whether the feasibility conclusion of the designated solution strategy meets the preset conditions.

[0133] It should be noted that, in the above content, this application implemented the Monte Carlo algorithm to evaluate the performance of the currently selected target solution strategy. The value function indicates that the target solution strategy has some drawbacks. Although the overall performance is satisfactory, the target solution strategy fails in some extreme or special cases. Moreover, sometimes there may be no associated candidate solution strategies for some unsolved events, requiring the summarization of a better strategy based on the changes between states to recommend to the user. Therefore, in an optional implementation, this application can also construct an optimal solution strategy for the unsolved event to recommend to the user, or optimize the existing candidate solution strategies and the already selected target solution strategy, as follows:

[0134] A policy can essentially be viewed as a set of actions taken between different states. By ensuring that all actions taken under different conditions are optimal, an optimal policy for resolving an event can be obtained. To find the optimal policy that maximizes the value function of each state variable, this application uses the Monte Carlo algorithm as a tool for iteratively evaluating and improving the policy until the value function converges to V* (maximum value). This process is called Monte Carlo control in a model-free environment. First, some basic theories and notations of the control process are introduced, and then they are applied to the process of selecting the optimal policy in this application.

[0135] The Monte Carlo control process consists of an iterative sequence of policy evaluation and policy improvement, such as... Figure 2C As shown, the general idea behind Monte Carlo control is Generalized Policy Iteration (GPI), which allows the policy evaluation and policy improvement processes to occur simultaneously. During the control process, a policy and its value function are defined, where the value function continuously changes to approximate the true value function of the current policy. The current policy is then repeatedly updated based on the value function. Each time a new policy is introduced, a new value function is used to approximate it, and the greedy policy is improved accordingly. This process continues until both the value function and the policy stabilize, meaning that the optimal policy has been found. Without knowing the transition probabilities, using the value function is insufficient to determine the best action to take in the next state, as shown in Equation 7 below:

[0136] Formula 7: π'(s)=argmaxa∈AR(a, s)+Pa(s, s`)V(s`)

[0137] Pa(s, s') represents the probability of obtaining state s to s' after taking action a, and R(a, s) represents the reward a for taking action in state s. Conversely, the action-value function shown in Equation 8 is introduced:

[0138] Formula 8: qπ(s,a): qπ(s,a)=Eπ[Gt|St=s,At=a]

[0139] Therefore, each state and its subsequent action are paired. When state s is visited and action a is taken, the corresponding reward is added to this pair. Policy evaluation is performed beforehand, with numerous experiments conducted and the average reward used to approximate the true action-value function for each state of the first visit method. qπ(s, a) is used, assuming no knowledge of the environment, as it is only necessary to find the pair of s where the action state will have the maximum action value and update the action to the current resolution policy used to resolve the event.

[0140] Furthermore, the general idea behind policy improvement is a greedy algorithm, which compares qπ(s, a) for each visited state, and improves the policy by updating all possible operations. This allows us to obtain the maximum operation value, i.e., π`(s) = argmaxa∈Aq(s, a). Here, when improving the policy, π and π` are any pair of deterministic policies such that for all s, qπ(s, π`(s)) ≥ vπ(s). The policy π` must be equal to or better than π, i.e., vπ`(s) ≥ vπ(s). This theorem applies to the greedy law because for all states s, qπ(s, π`(s)) = qπ(s, argmaxqπ(s, a)) = maxaqπ(s, a) ≥ qπ(s, π(s)) = vπ(s).

[0141] Therefore, this theorem guarantees that the algorithm is at least as good as the previous one, and the process will continuously improve, eventually converging to the optimal policy q*. As mentioned earlier, the action-value function does not require information about the environment. One problem with greedy policy improvement is that it can be too greedy. If a greedy policy is chosen decisively every time, the improvement may converge quickly, potentially missing opportunities to explore other possibilities and thus missing the optimal policy. Therefore, an ∈-greedy algorithm can be used to solve this problem. It introduces a small probability parameter ∈ and π(a|s), which is a probability distribution applied to action a when accessing state s. In each state, we choose the greedy policy ∈ with a probability of 1 - ∈ + ∈ / |A| and a random action with a smaller probability: ∈ / |A|, where |A| is the total number of operations in the operation space. For example, if ∈ = 0.5, a coin toss can be used to decide whether to adopt the greedy action or randomly choose another action, ensuring that the greedy policy is chosen in most cases without ignoring other possibilities. This method is called ∈-greedy policy improvement. It should be noted that a better approach is to continuously change ∈, which decreases with the number of games and eventually approaches 0 with the number of simulations. In other words, as the process unfolds, there is a higher chance of choosing a greedy strategy. Similarly, it can be proven that the ∈-greedy algorithm improves its strategy using the policy improvement theorem.

[0142] Therefore, using the above approach, this application constructs multiple state variables and employs the Monte Carlo algorithm to simulate the execution of the event to be resolved according to the possible values ​​of each state variable. During execution, a possible value is randomly selected for each state variable, and the selected value is used to assign values ​​to the corresponding state variables, resulting in multiple assigned state variables. Then, according to the change order of the multiple state variables, the first state variable and the next state variable are determined, simulating the transition from the first state variable to the next state variable. During the transition, simulated action values ​​are obtained, and the predicted transition probability from the first state variable to the next state variable is determined. Obtain the action value function constructed using the Monte Carlo algorithm. Input the simulated action value and the predicted transition probability into the action value function. Use the output value of the action value function as the simulated operation value. Continue to determine the next specified state variable from multiple state variables, and calculate the simulated operation value of the next state variable and the specified state variable, until multiple state variables are traversed, and end the current simulation. Obtain multiple simulated operation values. Take the simulation operation value with the largest value from the multiple simulated operation values ​​as the simulation reward value after the end of this simulation.

[0143] Subsequently, the target possible values ​​corresponding to each state variable in this simulation are obtained. Each state variable is labeled using the target possible values ​​corresponding to each state variable. The simulation reward value obtained after the simulation ends is obtained. The labeled multiple state variables and simulation reward values ​​are associated to obtain a simulation data set, which is then recorded.

[0144] Next, the Monte Carlo algorithm is re-emerged. The event to be resolved is simulated using the possible values ​​of each of the multiple state variables. New possible values ​​for each state variable are then acquired and labeled. The new simulation reward values ​​are associated with the newly labeled state variables to generate new simulation data sets. This process continues until the number of simulation rounds reaches a second preset number, resulting in a second preset number of simulation data sets. Finally, the target simulation data set with the highest simulation reward value is extracted from the second preset number of simulation data sets. The possible values ​​corresponding to the multiple state variables in the target simulation data set are combined to obtain a value set. This value set serves as the optimal solution strategy for the event to be resolved and is pushed to the user's terminal.

[0145] For example, define multiple state variables, each consisting of 13 parameters. The first parameter represents the current number of rolls, the next six parameters represent the number of each face held by the user, and the last six parameters, representing the number of each face, represent the number of rolls in the current round. For example, the state [2,1,0,1,0,0,0,0,1,3,0,2,0,4,0,5] indicates that the user holds an Ace and a Three-Point card, and played three Two-Point cards in the second round. Next, assuming all A(1), 1 is retained, for the initial policy, randomly select from six possible actions: retain one of 2, 3, 4, 5, or 6, or retain no points. The policy can be viewed as a set of actions corresponding to different states. That is, it's like a table where you can see what action to take when facing different states. Therefore, the number of simulations should be as large as possible to access all possible states, and the access time should be long enough. As shown in Table 3 below, the number of states does not change much when the simulation reaches 500,000. Therefore, this is considered the optimal number of trials.

[0146] Game rounds 500 5000 50,000 500,000 Number of states 1233 1788 2224 2338

[0147] Table 3

[0148] It should be noted that in practical applications, the above process can be implemented by building a tool and integrating it onto a platform, allowing the platform to perform calculations by calling the tool. When building this tool, similar to policy evaluation, each policy can be treated as a scenario, and a score can be assigned for different outcomes. Initially, the score for any policy is zero; the final score is the reward value for all visited states. Then, the action-value function is modified to approximate the true value under the current policy, and then improved. In practical applications, through iteration, the policy was optimized after 500,000 simulations. Specifically, as shown in Table 4 below, Table 4 displays information on the optimal action to take when a specific state is scrolled, and each occurrence of a point is preserved. For example, the first row indicates that points 1 and 6 should be kept in the state "1000000210002", meaning there are two points 1, one point 2, and two points 6 starting from the first scroll. Compared to the previously derived general policy, Table 4 is more specific because actions can be performed based on the state.

[0149] States Action 1000000210002 6 2000002011100 1 1000000000230 5 2010000012010 1 1000000111200 1 2100000011200 4 2100000100021 5

[0150] Table 4

[0151] It should be noted that for the dice game provided by the platform, the method proposed in this application will still be used if the game rules change. Suppose that 1 can no longer be considered any number; the number of dice changes, for example, rolling 4 dice; the scoring scheme changes, etc. In this case, the overall approach to solving these problems remains the same, only the variables and definitions in the model are modified. For example, in the first case, 1 is not retained by default; now there are 7 actions to take: retain 1, 2, 3, 4, 5, 6, or not retain any dice. Similar modifications are needed for the remaining cases. Thus, the composite strategy of this application performs better than a single strategy. Furthermore, in practical applications, Monte Carlo control can be introduced to simultaneously update candidate solutions and evaluate them to converge to the maximum value, resulting in better performance.

[0152] The method provided in this application embodiment enables multiple evaluations of the solution strategy, ensuring to a large extent that the solution strategy pushed to the user has the best effect, improving the accuracy of strategy recommendation, reducing errors, and truly solving problems for the user.

[0153] Furthermore, as Figure 1 In a specific implementation of the method, this application provides an event resolution strategy push device, such as... Figure 3 As shown, the device includes: an acquisition module 301, a scoring module 302, an evaluation module 303, and a push module 304.

[0154] The acquisition module 301 is used to respond to an event resolution request, determine the event to be resolved indicated by the event resolution request, and acquire multiple candidate resolution strategies associated with the event to be resolved;

[0155] The scoring module 302 is used to score each candidate solution strategy according to the expected effect value corresponding to each candidate solution strategy among the plurality of candidate solution strategies, and obtain the strategy score value corresponding to each candidate solution strategy.

[0156] The evaluation module 303 is used to extract the candidate solution strategy with the largest strategy score from the multiple candidate solution strategies as the target solution strategy, and to evaluate the feasibility of the target solution strategy using the Monte Carlo algorithm to obtain a feasibility conclusion of the target solution strategy.

[0157] The push module 304 is used to determine the user who initiated the event resolution request when the feasibility conclusion meets the preset conditions, and to push the target resolution strategy to the terminal held by the user.

[0158] In a specific application scenario, the acquisition module 301 is used to acquire the event identifier of the event to be resolved, query multiple optional result values ​​associated with the event identifier, and use each of the multiple optional result values ​​as a candidate solution strategy to obtain the multiple candidate solution strategies; or, determine a preset number of combinations, arrange and combine the multiple optional result values ​​according to the preset number of combinations to obtain multiple result value arrays, and use each of the multiple result value arrays as a candidate solution strategy to obtain the multiple candidate solution strategies, wherein the number of optional result values ​​included in each result value array is equal to the preset number of combinations.

[0159] In a specific application scenario, the scoring module 302 is used to read a first optional result value from each of the plurality of candidate solution strategies; determine the acquisition method of the first optional result value; query the acquisition probability of the first optional result value based on the acquisition method; acquire the expected effect value associated with the first optional result value; calculate the product of the expected effect value and the acquisition probability; and use the product as the strategy score value corresponding to the candidate solution strategy.

[0160] In a specific application scenario, the evaluation module 303 is used to read a second optional result value from the target solution strategy, and use the method of obtaining the second optional result value as the target acquisition method; simulate the execution of the target acquisition method, collect the actions that occur during the simulation, and determine the preset action scores associated with the actions that occur, and obtain the simulation result value obtained after the simulation execution is completed; obtain the expected reward function constructed using the Monte Carlo algorithm, input the action value of the actions that occur and the preset action scores into the expected reward function, and use the output value of the expected reward function as the expected reward value of this simulation process; associate the simulation result value with the expected reward value and use it as the simulation round result of this simulation process, and record the simulation round result; re-simulate the execution of the target acquisition method, and regenerate the simulation round result of the simulation process until the number of simulations reaches a first preset number, obtain the first preset number of simulation round results, and use the first preset number of simulation round results as the feasibility conclusion of the target solution strategy.

[0161] In a specific application scenario, the evaluation module 303 is further used to obtain the preset conditions, extract the preset conditions including standard result value, standard return value and quantity threshold; read a first preset number of simulation round results from the feasibility conclusion, count the number of results of the target simulation round, the target simulation round result being the simulation round result in the first preset number of simulation round results where the simulation result value is equal to the standard result value and the expected return value is greater than or equal to the standard return value; compare the number of results with the quantity threshold; accordingly, if the number of results is greater than or equal to the quantity threshold, then it is determined that the feasibility conclusion meets the preset conditions; wherein, if the number of results is less than the quantity threshold, then it is determined that the feasibility conclusion does not meet the preset conditions; determine the remaining candidate solution strategies other than the target solution strategy from the multiple candidate solution strategies; extract the candidate solution strategy with the largest strategy score from the remaining candidate solution strategies as the designated solution strategy; and use the Monte Carlo algorithm to evaluate the feasibility of the designated solution strategy to obtain the feasibility conclusion of the designated solution strategy, and determine whether the feasibility conclusion of the designated solution strategy meets the preset conditions.

[0162] In specific application scenarios, the device also includes:

[0163] A construction module is used to construct multiple state variables if there are no associated candidate resolution strategies for the event to be resolved.

[0164] The simulation module is used to simulate the execution of the event to be solved by employing the Monte Carlo algorithm according to the possible values ​​of each of the plurality of state variables;

[0165] The annotation module is used to obtain the target possible values ​​corresponding to each state variable in this simulation, and to annotate each state variable using the target possible values ​​corresponding to each state variable.

[0166] The association module is used to obtain the simulation reward value obtained after the simulation ends, associate the labeled multiple state variables with the simulation reward value to obtain a simulation data group, and record the simulation data group;

[0167] The simulation module is also used to re-emulate the Monte Carlo algorithm, simulate the execution of the event to be resolved according to the possible values ​​of each of the multiple state variables, and re-obtain and label the new possible values ​​of each state variable. The new simulation reward value obtained in this simulation is associated with the re-labeled multiple state variables to generate a new set of simulation data until the number of simulation rounds reaches the second preset number, and the second preset number of simulation data sets is obtained.

[0168] The extraction module is used to extract the target simulation data group with the largest simulation return value from the second preset number of simulation data groups, combine multiple possible values ​​corresponding to multiple state variables included in the target simulation data group to obtain a value group, and use the data group as the optimal solution strategy for the event to be solved.

[0169] The push module 304 is also used to push the optimal solution strategy to the user's terminal.

[0170] In a specific application scenario, this simulation module is used to: randomly select a possible value for each state variable based on the possible values ​​corresponding to each state variable; assign the selected possible value to the corresponding state variable to obtain the assigned multiple state variables; determine the first state variable and the next state variable from the multiple state variables according to the change order corresponding to the multiple state variables; simulate the transition from the first state variable to the next state variable, obtain simulated action values ​​during the transition, and determine the predicted transition probability from the first state variable to the next state variable; obtain an action value function constructed using the Monte Carlo algorithm, input the simulated action value and the predicted transition probability into the action value function, and use the output value of the action value function as the simulated operation value; continue to determine the next specified state variable from the multiple state variables, and calculate the simulated operation values ​​of the next state variable and the specified state variable, until the multiple state variables are traversed, ending the current simulation and obtaining multiple simulated operation values; and use the simulation operation value with the largest value among the multiple simulated operation values ​​as the simulation reward value after the end of the current simulation.

[0171] The apparatus provided in this application embodiment, in response to an event resolution request, determines the event to be resolved indicated by the event resolution request, obtains multiple candidate resolution strategies associated with the event to be resolved, scores each candidate resolution strategy according to the expected effect value corresponding to each candidate resolution strategy, obtains a strategy score value corresponding to each candidate resolution strategy, extracts the candidate resolution strategy with the largest strategy score value from the multiple candidate resolution strategies as the target resolution strategy, and uses a Monte Carlo algorithm to evaluate the feasibility of the target resolution strategy to obtain a feasibility conclusion of the target resolution strategy. When the feasibility conclusion meets preset conditions, it determines the user who initiated the event resolution request and pushes the target resolution strategy to the user's terminal, realizing multiple evaluation of the resolution strategy, ensuring to a large extent that the resolution strategy pushed to the user has the best effect, improving the accuracy of strategy recommendation, reducing errors, and truly solving problems for users.

[0172] It should be noted that other corresponding descriptions of the functional units involved in the event resolution strategy push device provided in this application embodiment can be found in the following references. Figure 1 and Figures 2A to 2C The corresponding description in [the document] will not be repeated here.

[0173] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0174] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0175] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

[0176] In an exemplary embodiment, see Figure 4 The invention also provides a computer device including a bus, a processor, a memory, and a communication interface. It may also include an input / output interface and a display device, wherein the various functional units can communicate with each other via the bus. The memory stores a computer program, and the processor executes the program stored in the memory to perform the event resolution strategy push method described in the above embodiments.

[0177] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the push method for the event resolution strategy.

[0178] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented in hardware or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0179] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application.

[0180] Those skilled in the art will understand that the modules in the apparatus of the implementation scenario can be distributed within the apparatus of the implementation scenario as described, or they can be located in one or more apparatuses different from this implementation scenario, with corresponding changes. The modules of the above-described implementation scenario can be combined into one module, or they can be further divided into multiple sub-modules.

[0181] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of the implementation scenario.

[0182] The above disclosures are only a few specific implementation scenarios of this application. However, this application is not limited to these. Any variations that can be conceived by those skilled in the art should fall within the protection scope of this application.

Claims

1. A method for pushing event resolution strategies, characterized in that, include: In response to an event resolution request, determine the event to be resolved indicated by the event resolution request, and obtain multiple candidate resolution strategies associated with the event to be resolved; Based on the expected effect value corresponding to each of the plurality of candidate solution strategies, each candidate solution strategy is scored to obtain a strategy score value corresponding to each candidate solution strategy; wherein, for each of the plurality of candidate solution strategies, a first optional result value is read from the candidate solution strategy; the acquisition method of the first optional result value is determined, and the acquisition probability of the first optional result value based on the acquisition method is queried; the expected effect value associated with the first optional result value is obtained, the product of the expected effect value and the acquisition probability is calculated, and the product is used as the strategy score value corresponding to the candidate solution strategy; The candidate solution strategy with the largest strategy score is selected as the target solution strategy from the multiple candidate solution strategies, and the feasibility of the target solution strategy is evaluated using the Monte Carlo algorithm to obtain a feasibility conclusion for the target solution strategy. When the feasibility conclusion meets the preset conditions, the user who initiated the event resolution request is identified, and the target resolution strategy is pushed to the user's terminal.

2. The method according to claim 1, characterized in that, The step of obtaining multiple candidate resolution strategies associated with the event to be resolved includes: Obtain the event identifier of the event to be resolved, query multiple optional result values ​​associated with the event identifier, and use each of the multiple optional result values ​​as a candidate resolution strategy to obtain the multiple candidate resolution strategies; or, A preset number of combinations is determined, and the multiple optional result values ​​are arranged and combined according to the preset number of combinations to obtain multiple result value arrays. Each result value array in the multiple result value arrays is used as a candidate solution strategy to obtain the multiple candidate solution strategies. The number of optional result values ​​included in each result value array is equal to the preset number of combinations.

3. The method according to claim 1, characterized in that, The evaluation of the feasibility of the objective solution strategy using the Monte Carlo algorithm, to obtain a feasibility conclusion for the objective solution strategy, includes: The second optional result value read in the target solution strategy is used as the target acquisition method; The target acquisition method is simulated, and during the simulation, the actions that occur are collected, and the preset action scores associated with the actions that occur are determined. The simulation result value is obtained after the simulation is completed. Obtain the expected reward function constructed using the Monte Carlo algorithm, input the action value of the action and the preset action score into the expected reward function, and use the output value of the expected reward function as the expected reward value for this simulation process; The simulation result value is correlated with the expected return value and used as the simulation round result of this simulation process, and the simulation round result is recorded; The target acquisition method is re-simulated and the simulation round results are regenerated until the number of simulations reaches a first preset number. The first preset number of simulation round results are obtained, and the first preset number of simulation round results are used as the feasibility conclusion of the target solution strategy.

4. The method according to claim 1, characterized in that, After selecting the candidate solution strategy with the highest strategy score from the multiple candidate solution strategies as the target solution strategy, and evaluating the feasibility of the target solution strategy using the Monte Carlo algorithm to obtain a feasibility conclusion for the target solution strategy, the method further includes: The preset conditions are obtained, and the preset conditions include standard result values, standard return values, and quantity thresholds; The feasibility conclusion reads a first preset number of simulation round results, and counts the number of results for the target simulation round. The target simulation round results are the simulation round results in the first preset number of simulation round results where the simulation result value is equal to the standard result value and the expected return value is greater than or equal to the standard return value. The number of results is compared with the number threshold; Accordingly, if the number of results is greater than or equal to the number threshold, then the feasibility conclusion is determined to meet the preset condition; If the number of results is less than the number threshold, it is determined that the feasibility conclusion does not meet the preset condition. Among the multiple candidate solution strategies, the remaining candidate solution strategies other than the target solution strategy are determined. Among the remaining candidate solution strategies, the candidate solution strategy with the largest strategy score is extracted as the designated solution strategy. The feasibility of the designated solution strategy is evaluated using the Monte Carlo algorithm to obtain the feasibility conclusion of the designated solution strategy, and it is determined whether the feasibility conclusion of the designated solution strategy meets the preset condition.

5. The method according to claim 1, characterized in that, The method further includes: If there are no associated candidate solutions for the event to be resolved, then multiple state variables are constructed; The Monte Carlo algorithm is used to simulate the execution of the event to be resolved according to the possible values ​​of each of the plurality of state variables; Obtain the target possible values ​​corresponding to each state variable in this simulation, and label each state variable using the target possible values ​​corresponding to each state variable; Obtain the simulation reward value after the simulation ends, associate the labeled multiple state variables with the simulation reward value to obtain a simulation data set, and record the simulation data set; The Monte Carlo algorithm is re-adopted, and the event to be resolved is simulated and executed according to the possible values ​​of each of the multiple state variables. The new possible values ​​of the target corresponding to each state variable are obtained and labeled. The new simulation reward value obtained in this simulation is associated with the relabeled multiple state variables to generate a new simulation data set. The simulation is repeated until the number of rounds reaches the second preset number, and the second preset number of simulation data sets are obtained. Extract the target simulation data group with the largest simulated return value from the second preset number of simulation data groups, combine the multiple possible values ​​corresponding to the multiple state variables included in the target simulation data group to obtain a value group, and use the data group as the optimal solution strategy for the event to be solved. The optimal solution strategy is pushed to the user's terminal.

6. The method according to claim 5, characterized in that, The step of simulating the event to be resolved using the Monte Carlo algorithm, according to the possible values ​​of each of the plurality of state variables, includes: Based on the possible values ​​corresponding to each state variable, a possible value is randomly selected for each state variable, and the selected possible value is used to assign a value to the corresponding state variable to obtain the multiple state variables after assignment. According to the order of change of the multiple state variables, determine the first state variable and the next state variable of the first state variable among the multiple state variables; Simulate the process of transitioning from the first state variable to the next state variable, acquire simulated action values ​​during the transition, and determine the predicted transition probability of the transition from the first state variable to the next state variable; Obtain the action value function constructed using the Monte Carlo algorithm, input the simulated action value and the predicted transition probability into the action value function, and use the output value of the action value function as the simulated operation value; Continue to determine the next specified state variable of the next state variable among the multiple state variables, and calculate the simulation operation value of the next state variable and the specified state variable, until the multiple state variables are traversed, and end the current simulation, and obtain multiple simulation operation values; The simulation operation value with the largest value among the multiple simulation operation values ​​is taken as the simulation reward value after the end of this simulation.

7. A device for pushing out an event resolution strategy, characterized in that, include: The acquisition module is used to respond to an event resolution request, determine the event to be resolved indicated by the event resolution request, and acquire multiple candidate resolution strategies associated with the event to be resolved; The scoring module is used to score each candidate solution strategy based on the expected effect value corresponding to each candidate solution strategy among the plurality of candidate solution strategies, and obtain a strategy score value corresponding to each candidate solution strategy; wherein, the scoring module is used to, for each candidate solution strategy among the plurality of candidate solution strategies, read a first optional result value from the candidate solution strategy; determine the acquisition method of the first optional result value, query the acquisition probability of the first optional result value based on the acquisition method; obtain the expected effect value associated with the first optional result value, calculate the product of the expected effect value and the acquisition probability, and use the product as the strategy score value corresponding to the candidate solution strategy; The evaluation module is used to extract the candidate solution strategy with the largest strategy score from the multiple candidate solution strategies as the target solution strategy, and to evaluate the feasibility of the target solution strategy using the Monte Carlo algorithm to obtain a feasibility conclusion of the target solution strategy. The push module is used to determine the user who initiated the event resolution request when the feasibility conclusion meets the preset conditions, and to push the target resolution strategy to the terminal held by the user.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.