Resource allocation strategy adjustment method and device, computer equipment and storage medium
By obtaining and adjusting the original strategy and feedback data of the resource allocation platform, and using the policy adjustment system and model for multiple adjustments and convergence, the problem of low accuracy in resource allocation strategy adjustment is solved, and more accurate and efficient resource allocation is achieved.
Patent Information
- Application Number
- CN202311778986.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, the accuracy of resource allocation strategy adjustment is low, especially the lack of reference strategies in the early stage of machine learning model training, resulting in low early accuracy.
A resource allocation strategy adjustment method is proposed. By obtaining the original policy and feedback data of the target resource allocation platform, using the policy adjustment system and model for multiple adjustments and convergence, to generate a more accurate resource allocation strategy.
Through multiple adjustments and fusions, the accuracy and effectiveness of resource allocation strategies are improved, and the amount of sample data required for manual participation and model training is reduced.
Smart Images

Figure CN120197844A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of fintech, and in particular, to a method and device for adjusting a resource allocation strategy, a computer device, and a storage medium. Background Art
[0002] Currently, the biggest challenge in the field of resource allocation lies in how to adjust the optimal resource allocation strategy. In related technologies, the adjustment of the resource allocation strategy is mainly completed by a machine learning model. However, with the endless emergence of problems in resource allocation, the accuracy of the resource allocation strategy adjusted by the machine learning model decreases, and there is no reference resource allocation strategy in the initial stage of training the machine learning model. Therefore, the accuracy of the machine learning model in the early stage of resource allocation strategy adjustment is relatively low. Therefore, how to make the adjustment of the resource allocation strategy in different stages more accurate has become an urgent technical problem to be solved. Summary of the Invention
[0003] The main purpose of the embodiments of the present application is to propose a method and device for adjusting a resource allocation strategy, a computer device, and a storage medium, aiming to make the adjustment of the resource allocation strategy in different stages more accurate.
[0004] To achieve the above object, a first aspect of the embodiments of the present application proposes a method for adjusting a resource allocation strategy, the method including:
[0005] Obtain the original resource allocation strategy of the target resource allocation platform;
[0006] Based on the original resource allocation strategy, obtain the feedback data of the target object to obtain the original feedback data; wherein, the original feedback data represents the feedback situation of the target object on the target resource allocation platform based on the original resource allocation strategy;
[0007] Send the original resource allocation strategy and the original feedback data to a preset policy adjustment system, and receive the first resource allocation strategy feedback by the policy adjustment system according to the original resource allocation strategy and the original feedback data;
[0008] According to a preset policy adjustment model and the original feedback data, perform policy adjustment on the original resource allocation strategy to obtain a second resource allocation strategy;
[0009] Obtain the original usage parameters output by the policy adjustment system, and perform fusion processing on the first resource allocation strategy and the second resource allocation strategy according to the original usage parameters to obtain a preliminary resource allocation strategy; wherein, the original usage parameters represent the usage probability of the first resource allocation strategy;
[0010] Obtain the feedback data of the target object based on the preliminary resource allocation strategy to obtain preliminary feedback data; wherein, the preliminary feedback data characterizes the feedback of the target object on the target resource allocation platform based on the preliminary resource allocation strategy.
[0011] Adjust the preliminary resource allocation strategy according to the preliminary feedback data.
[0012] In some embodiments, the adjusting the preliminary resource allocation strategy according to the preliminary feedback data includes:
[0013] Adjust the original usage parameters according to the preliminary feedback data to obtain target usage parameters.
[0014] Send the preliminary resource allocation strategy and the preliminary feedback data to the strategy adjustment system, and receive the third resource allocation strategy feedback by the strategy adjustment system according to the preliminary resource allocation strategy and the preliminary feedback data.
[0015] Perform strategy adjustment on the preliminary resource allocation strategy through the strategy adjustment model, the target usage parameters and the preliminary feedback data to obtain a fourth resource allocation strategy.
[0016] Adjust the preliminary resource allocation strategy according to the target usage parameters, the third resource allocation strategy and the fourth resource allocation strategy.
[0017] In some embodiments, the preliminary feedback data includes: platform usage growth data and platform resource loss data; the adjusting the original usage parameters according to the preliminary feedback data to obtain target usage parameters includes:
[0018] Perform standardization processing on the platform usage growth data to obtain platform usage growth parameters.
[0019] Perform standardization processing on the resource loss data to obtain resource reduction parameters.
[0020] Adjust the original usage parameters according to the platform usage growth parameters and the resource reduction parameters to obtain the target usage parameters.
[0021] In some embodiments, the performing strategy adjustment on the preliminary resource allocation strategy and the preliminary feedback data through the strategy adjustment model and the target usage parameters to obtain a fourth resource allocation strategy includes:
[0022] Obtain the model loss data of the strategy adjustment model according to the target usage parameters.
[0023] Adjust the parameters of the policy adjustment model according to the model loss data;
[0024] Use the adjusted policy adjustment model to adjust the preliminary resource allocation policy and the preliminary feedback data to obtain a fourth resource allocation policy.
[0025] In some embodiments, obtaining the model loss data of the policy adjustment model according to the target usage parameters includes:
[0026] Compare the target usage parameters with a preset threshold to obtain a comparison result;
[0027] If the comparison result indicates that the target usage parameters are greater than the preset threshold, use the loss data between the first resource allocation policy and the second resource allocation policy as the model loss data;
[0028] If the comparison result indicates that the target usage parameters are less than or equal to the preset threshold, use the preliminary feedback data as the model loss data.
[0029] In some embodiments, the policy adjustment model includes at least one policy adjustment network; using the adjusted policy adjustment model and the preliminary feedback data to adjust the preliminary resource allocation policy to obtain a fourth resource allocation policy includes:
[0030] Input the original resource allocation policy and the original feedback data into each adjusted policy adjustment network for policy adjustment to obtain at least one candidate resource allocation policy;
[0031] Fuse at least one of the candidate resource allocation policies to obtain the fourth resource allocation policy.
[0032] In some embodiments, obtaining the original usage parameters output by the policy adjustment system and fusing the first resource allocation policy and the second resource allocation policy according to the original usage parameters to obtain a preliminary resource allocation policy includes:
[0033] Obtain the original usage parameters output by the policy adjustment system;
[0034] Extract first resource allocation parameters from the first resource allocation policy;
[0035] Extract second resource allocation parameters from the second resource allocation policy;
[0036] Select a target parameter adjustment operation from preset candidate parameter adjustment operations according to the first resource allocation parameters and the second resource allocation parameters;
[0037] Perform the target parameter adjustment operation on the first resource allocation parameter and the second resource allocation parameter according to the original usage parameter to obtain a target resource allocation parameter;
[0038] Construct the preliminary resource allocation strategy according to the target resource allocation parameter.
[0039] To achieve the above object, a second aspect of the embodiments of the present application proposes a resource allocation strategy adjustment device, and the device includes:
[0040] A strategy acquisition module, configured to acquire an original resource allocation strategy of a target resource allocation platform;
[0041] An original feedback acquisition module, configured to acquire feedback data of a target object based on the original resource allocation strategy to obtain original feedback data; wherein, the original feedback data characterizes the feedback situation of the target object on the target resource allocation platform based on the original resource allocation strategy;
[0042] A sending module, configured to send the original resource allocation strategy and the original feedback data to a preset strategy adjustment system, and receive a first resource allocation strategy fed back by the strategy adjustment system according to the original resource allocation strategy and the original feedback data;
[0043] A model adjustment module, configured to perform strategy adjustment on the original resource allocation strategy according to a preset strategy adjustment model and the original feedback data to obtain a second resource allocation strategy;
[0044] A parameter acquisition module, configured to acquire an original usage parameter output by the strategy adjustment system, and perform fusion processing on the first resource allocation strategy and the second resource allocation strategy according to the original usage parameter to obtain a preliminary resource allocation strategy; wherein, the original usage parameter characterizes the usage probability of the first resource allocation strategy;
[0045] A preliminary feedback acquisition module, configured to acquire feedback data of the target object based on the preliminary resource allocation strategy to obtain preliminary feedback data; wherein, the preliminary feedback data characterizes the feedback situation of the target object on the target resource allocation platform based on the preliminary resource allocation strategy;
[0046] A strategy adjustment module, configured to adjust the preliminary resource allocation strategy according to the preliminary feedback data.
[0047] To achieve the above object, a third aspect of the embodiments of the present application proposes a computer device, and the computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0048] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method described in the first aspect above.
[0049] The resource allocation strategy adjustment method, device, computer equipment and storage medium provided by the present application integrate a strategy adjustment system and a strategy adjustment model to update the resource allocation strategy based on the feedback data of the target object, and update the usage parameters after each resource allocation strategy adjustment is completed. Therefore, after the resource allocation strategy is output by the strategy adjustment model and the strategy adjustment system, the two resource allocation strategies are combined according to the usage parameters, which is convenient to combine the resource allocation strategies with different usage parameters at different stages to output a more optimal and accurate resource allocation strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of the resource allocation strategy adjustment method provided by the embodiments of the present application;
[0051] Figure 2 is a schematic structural diagram of a tree classifier in the resource allocation strategy adjustment method provided by the embodiments of the present application;
[0052] Figure 3 is Figure 1 a flowchart of step S105 in
[0053] Figure 4 is Figure 1 a flowchart of step S107 in
[0054] Figure 5 is Figure 4 a flowchart of step S401 in
[0055] Figure 6 is Figure 4 a flowchart of step S403 in
[0056] Figure 7 is Figure 6 a flowchart of step S601 in
[0057] Figure 8 is a schematic structural diagram of a strategy adjustment model in the resource allocation strategy adjustment method provided by the embodiments of the present application;
[0058] Figure 9 is Figure 6 a flowchart of step S603 in
[0059] Figure 10 is a detailed flowchart of the resource allocation strategy adjustment method provided by the embodiments of the present application;
[0060] Figure 11 It is a schematic structural diagram of a resource allocation strategy adjustment device provided by an embodiment of the present application;
[0061] Figure 12 It is a schematic hardware structure diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.
[0063] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0065] First, several terms involved in the present application are analyzed:
[0066] Reinforcement Learning (RL): It is a machine learning method aimed at learning how to make optimal decisions by interacting with the environment. In reinforcement learning, a model (referred to as an agent) interacts with the environment by observing the state of the environment and performing actions. Based on the feedback from the environment, the agent can obtain rewards or punishments to guide its learning process.
[0067] Linear regression network: It is a common predictive model used in statistics and machine learning, which attempts to predict continuous outputs by modeling the linear relationship between independent variables (features) and dependent variables (targets).
[0068] Boosted tree network: It is an ensemble learning method that forms a powerful model by combining multiple weak classifiers (i.e., decision trees). Boosting refers to training a series of weak classifiers serially, and each classifier weights and corrects the errors of the previous classifier, thereby gradually improving the accuracy of the overall model.
[0069] Logistic regression network: A statistical model for classification problems, which is a generalized linear model that transforms a linear relationship into a probability by mapping the output of a linear equation to the range of a sigmoid function (also known as the logistic function). The logistic regression model can be used for binary classification problems and can also be extended to multi-classification problems.
[0070] Agent: Refers to an artificial intelligence system with the abilities of perception, thinking, and action. It can obtain environmental information through sensors (such as cameras, microphones, etc.), process and analyze data through algorithms and models, and then use a decision-making system to take actions.
[0071] In the platform of fintech, a numerical system will be set up, such as a game currency mechanism, a points system, and platform discounts, etc. Therefore, each numerical system involves a resource allocation strategy. How to find a resource allocation strategy that suits each platform is a long-term optimization process and faces several challenges. In the process of adjusting the resource allocation strategy, not only the issue of resource consumption needs to be considered, but also the operation effect of the platform needs to be improved. In the process of adjusting the resource allocation strategy, if only relying on professional operation personnel and expert experience, it will consume a large amount of manpower. Only using a model to achieve intelligent adjustment of the resource allocation strategy, because a large amount of sample data is also required in the early stage of model training, it is difficult to construct the sample data, and a small amount of sample data is difficult to enable the model to achieve accurate adjustment of the strategy.
[0072] Based on this, the embodiments of this application provide a method and device for adjusting a resource allocation strategy, a computer device, and a storage medium, aiming to jointly adjust the strategy by setting up a strategy adjustment system and a strategy adjustment model. In the process of strategy adjustment, the resource allocation strategy will be adjusted based on the feedback data after the execution of the resource allocation strategy as a reference to adjust a better resource allocation strategy. At the same time, when combining the resource allocation strategies output by the strategy adjustment model and the strategy adjustment system, the two resource allocation strategies are combined according to the usage parameters, so as to use different usage parameters at different stages to combine the resource allocation strategies to output a better and more accurate resource allocation strategy.
[0073] The method and device for adjusting a resource allocation strategy, a computer device, and a storage medium provided by the embodiments of this application are specifically described through the following embodiments. First, the method for adjusting a resource allocation strategy in the embodiments of this application is described.
[0074] The embodiments of this application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0075] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0076] The resource allocation strategy adjustment method provided by the embodiments of this application relates to the field of artificial intelligence technology. The resource allocation strategy adjustment method provided by the embodiments of this application can be applied to terminals, can also be applied to server sides, and can also be software running on terminals or server sides. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the resource allocation strategy adjustment method, etc., but is not limited to the above forms.
[0077] This application can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0078] Figure 1 is an optional flowchart of the resource allocation strategy adjustment method provided by the embodiments of this application, Figure 1 The method in may include but is not limited to steps S101 to S107.
[0079] Step S101, obtain the original resource allocation strategy of the target resource allocation platform;
[0080] Step S102: Obtain the feedback data of the target object based on the original resource allocation strategy to get the original feedback data. The original feedback data characterizes the feedback of the target object on the target resource allocation platform based on the original resource allocation strategy.
[0081] Step S103: Send the original resource allocation strategy and the original feedback data to a preset policy adjustment system, and receive the first resource allocation strategy fed back by the policy adjustment system according to the original resource allocation strategy and the original feedback data.
[0082] Step S104: Adjust the original resource allocation strategy and the original feedback data according to a preset policy adjustment model and the original feedback data to obtain the second resource allocation strategy.
[0083] Step S105: Obtain the original usage parameters output by the policy adjustment system, and perform a fusion process on the first resource allocation strategy and the second resource allocation strategy according to the original usage parameters to obtain the preliminary resource allocation strategy. The original usage parameters characterize the usage probability of the first resource allocation strategy.
[0084] Step S106: Obtain the feedback data of the target object based on the preliminary resource allocation strategy to get the preliminary feedback data. The preliminary feedback data characterizes the feedback of the target object on the target resource allocation platform based on the preliminary resource allocation strategy.
[0085] Step S107: Adjust the preliminary resource allocation strategy according to the preliminary feedback data.
[0086] Steps S101 to S107 illustrated in the embodiments of the present application are as follows: By obtaining the original resource allocation strategy of the target resource allocation platform and the original feedback data after the target object executes the original resource allocation strategy, and then inputting the original feedback data and the original resource allocation strategy into the policy adjustment system and the policy adjustment model respectively. The original resource allocation strategy is adjusted to obtain the first resource allocation strategy through the policy adjustment system and the original feedback data, and the original resource allocation strategy is adjusted to obtain the second resource allocation strategy through the policy adjustment model and the original feedback data. After jointly adjusting the policy through the policy adjustment system and the policy adjustment model, in order to construct a more accurate resource allocation strategy, the first resource allocation strategy and the second resource allocation strategy are combined into a preliminary resource allocation strategy according to the original usage parameters. After the preliminary resource allocation strategy is constructed, the preliminary feedback data fed back by the target object is continuously obtained based on the preliminary resource allocation strategy, and then the preliminary resource allocation strategy is continuously adjusted according to the preliminary feedback data. Therefore, the embodiments of the present application integrate the policy adjustment system and the policy adjustment model to update the resource allocation strategy with the feedback data of the target object, no longer relying solely on manual work, reducing the burden on manual labor, constructing a better resource allocation strategy, and improving the operation effect of the target resource allocation platform through the updated resource allocation strategy.
[0087] In step S101 of some embodiments, the original resource allocation strategy is the strategy used by the target resource allocation platform for resource allocation, so that the target resource allocation platform can achieve resource allocation according to the original resource allocation strategy. It should be noted that in different fields, the target resource allocation platform and the original resource allocation strategy are different. For example, in the abnormal repair strategy, the target resource allocation platform is the repair resource allocation platform on the terminal, and the original resource allocation strategy is the repair resource allocation strategy for abnormal repair, that is, what repair resources to select to solve the abnormal problem. Among them, the repair resources can be abnormal repair tools, abnormal repair algorithms, etc. In the logistics field, the target resource allocation platform is the logistics platform, which is used for order resource and transportation resource allocation in the logistics process, and the original resource allocation strategy can be the order resource allocation strategy or the transportation resource allocation strategy. In the object management field, the original resource allocation platform is the integral allocation platform that allocates points to objects, and the original resource allocation strategy is the integral allocation strategy. Therefore, in different fields, the original resource allocation platform and the original resource allocation strategy are different, both representing a kind of resource allocation.
[0088] In an application scenario, taking the field of object management as an example, the object is a logistics courier, the original resource allocation strategy is an integral allocation strategy, and the integral allocation strategy can be characterized by the integral reward intensity of different scoring indicators and the uses of the integral. Among them, the uses of the integral can be participating in activities, redeeming goods, exchanging rights and interests, etc. It should be noted that for a lottery activity, the integral reward intensity is a rate of return. The price of redeeming goods is stable, and the exchanged rights and interests are the preferential intensity of goods, exemption from customer complaints, etc., and the total exchanged rights and interests can be adjusted. Therefore, taking the coefficients corresponding to the integral reward intensity and the integral uses as adjustable parameters of the integral allocation platform, adjusting the resource allocation strategy, that is, adjusting the adjustable parameters, can improve the enthusiasm of the logistics courier and thus improve the operation effect of the integral allocation platform.
[0089] In step S102 of some embodiments, after the target resource allocation platform executes the original resource allocation strategy, it is necessary to analyze the operation effect of the target resource allocation platform, and the operation effect needs to be feedback by the target object. It should be noted that the original feedback data represents the feedback of the target object on the target resource allocation platform based on the original resource allocation strategy, and the feedback can be characterized by the satisfaction of the target object, check-in data, task completion rate, resource loss data, etc. In this embodiment, a data monitoring module is set up to monitor the data feedback by the target object in real time through the data monitoring module.
[0090] For example, in the field of object management, the original feedback data is the satisfaction, check-in volume and daily work feedback data of the logistics courier, etc. The daily work feedback data can be abnormal fluctuation warnings, regional imbalance warnings, courier cheating warnings, etc. At the same time, the original feedback data also includes the costs brought by the integral use, so as to judge whether it is necessary to adjust the original resource allocation strategy according to the original feedback data.
[0091] In step S103 of some embodiments, the strategy adjustment system is docked with multiple terminals, and each terminal is controlled by an operation object with experience in strategy adjustment in this field. So when the resource allocation strategy output by the strategy adjustment model is not yet in line with the operation requirements of the target resource allocation platform during the generation of the resource allocation strategy in the early stage, the resource allocation strategy output by the strategy adjustment system with experience in strategy adjustment is taken as the main one, so as to rely on the operation object with experience in resource adjustment to smooth the cold start of the strategy adjustment model. In addition, the strategy adjustment system can be a linear tree classifier, which is constructed by an object with sufficient experience in the field of resource allocation strategy adjustment. Therefore, inputting the original feedback data and the original resource allocation strategy into the tree classifier to output the first resource allocation strategy makes the construction of the first resource allocation strategy simple.
[0092] Such as Figure 2As shown, different resource allocation strategies with different directions are set on the tree classifier according to different feedback data. For example, if the original resource feedback data is participation and the original resource allocation strategy is the intensity of the integral activity, then the flow direction on the tree classifier can be that if the participation drops by n%, the activity intensity increases by 0.5 * n%, otherwise it increases by 0.1 * n%; where n is the preset threshold range of the participation drop. If the participation drop rate reaches n%, the activity intensity is increased according to a higher adjustment step, otherwise it is increased according to a smaller adjustment step.
[0093] In step S104 of some embodiments, in order to save manpower or improve the real-time performance of policy adjustment, a policy adjustment model is set, and the policy adjustment model is a machine learning model. It should be noted that the model structure is updated during the policy adjustment process to construct a policy adjustment model that can accurately perform policy adjustment. The original resource allocation policy is adjusted by the policy adjustment model and the original feedback data to obtain the second resource allocation policy, realizing automated policy adjustment and reducing the manpower of manual policy adjustment.
[0094] Please refer to Figure 3 , in some embodiments, step S105 may include but is not limited to steps S301 to S306:
[0095] Step S301, obtaining the original usage parameters output by the policy adjustment system;
[0096] Step S302, extracting the first resource allocation parameters from the first resource allocation policy;
[0097] Step S303, extracting the second resource allocation parameters from the second resource allocation policy;
[0098] Step S304, screening out the target parameter adjustment operation from the preset candidate parameter adjustment operations according to the first resource allocation parameter and the second resource allocation parameter;
[0099] Step S305, performing the target parameter adjustment operation on the first resource allocation parameter and the second resource allocation parameter according to the original usage parameters to obtain the target resource allocation parameters;
[0100] Step S306, constructing a preliminary resource allocation policy according to the target resource allocation parameters.
[0101] In step S301 of some embodiments, the original usage coefficient is the probability of adopting the first resource allocation policy, which can be defined as the usage probability of the resource allocation policy feedback by experienced operation objects, or can also be understood as the probability of adopting the manual opinion on resource allocation. It should be noted that the original usage parameter is α 原 , and the usage probability of the second resource allocation policy is β 原 , and β原 = 1 - α 原 Therefore, the original usage parameter characterizes the probability of using an artificial resource allocation strategy and also the probability of using the resource allocation strategy output by the strategy adjustment model. In this embodiment, if the original usage parameter is in the early stage of training the strategy adjustment model, α 原 = 1, that is, the resource allocation strategy output by the strategy adjustment model is not used. As the number of training times of the strategy adjustment model increases, the usage parameter gradually decreases to reduce manual participation and save manpower.
[0102] In steps S302 to S303 of some embodiments, since the resource allocation strategy involves the adjustment of multiple resource allocation parameters, it is necessary to extract the resource allocation parameters of each resource allocation strategy and determine the fusion method of the first resource allocation and the second resource allocation parameters. It should be noted that in the integral allocation platform, the resource allocation parameters can be the activity intensity, the commodity discount intensity, and so on.
[0103] In step S304 of some embodiments, the candidate parameter adjustment operation is a fusion method of the first resource allocation parameter and the second resource allocation parameter, and the candidate parameter adjustment operation can include weighted summation, summation, up and down adjustment, average value, and so on. Therefore, after determining the first resource allocation parameter and the second resource allocation parameter, a target parameter adjustment operation is selected from multiple candidate parameter adjustment operations to determine the fusion method of the first resource allocation parameter and the second resource allocation parameter.
[0104] In step S305 of some embodiments, the original usage parameter, the first resource allocation parameter, and the second resource allocation parameter are adjusted by the target parameter adjustment operation to obtain the target resource allocation parameter. For example, if the first resource allocation parameter and the second resource allocation parameter are the activity intensity and the target parameter adjustment operation is determined to be weighted summation, then the target resource allocation parameter is determined as shown in Equation (1):
[0105] α 目标 = α 原 * L1 + (1 - α 原 ) * L2 (1)
[0106] In the formula, L1 is the first resource allocation parameter and L2 is the second resource allocation parameter.
[0107] In step S306 of some embodiments, at least one target resource allocation parameter is integrated into a preliminary resource allocation strategy to make the construction of the resource allocation strategy more in line with the operation requirements of the target resource allocation platform.
[0108] In steps S301 to S306 illustrated in this embodiment, by selecting resource allocation parameters in each resource allocation strategy, then determining adjustment operations for each resource allocation parameter, different parameter adjustment operations for different resource allocation parameters are fused into target resource allocation parameters, and finally, based on at least one target resource allocation parameter, a preliminary resource allocation strategy is combined to construct a resource allocation strategy that better meets the operation requirements of the target resource allocation platform.
[0109] In step S106 of some embodiments, after completing an adjustment of the resource allocation strategy once, it is still necessary to obtain feedback data of the target object on the preliminary resource allocation strategy to obtain preliminary feedback data. It should be noted that the preliminary feedback data represents the resource consumption, the number of access times, and the satisfaction of the target object on the target resource allocation platform during the execution of the preliminary resource feedback strategy. The impact of the preliminary resource allocation strategy on the operation of the target resource allocation platform is judged by obtaining the preliminary feedback data.
[0110] Please refer to Figure 4 , in some embodiments, step S107 may include but is not limited to steps S401 to S404:
[0111] Step S401, adjust the original usage parameters according to the preliminary feedback data to obtain target usage parameters;
[0112] Step S402, send the preliminary resource allocation strategy and the preliminary feedback data to the strategy adjustment system, and receive the third resource allocation strategy feedback by the strategy adjustment system according to the preliminary resource allocation strategy and the preliminary feedback data;
[0113] Step S403, perform strategy adjustment on the preliminary resource allocation strategy through the strategy adjustment model, the target usage parameters, and the preliminary feedback data to obtain the fourth resource allocation strategy;
[0114] Step S404, adjust the preliminary resource allocation strategy according to the target usage parameters, the third resource allocation strategy, and the fourth resource allocation strategy.
[0115] In step S401 of some embodiments, it is determined whether to increase or decrease the original usage parameters according to the preliminary feedback data to obtain the target usage parameters. It should be noted that during the adjustment process of the original usage parameters, the loss data between the first resource allocation strategy and the second resource allocation strategy is also considered to adjust the original usage parameters through the loss data. Specifically, if the preliminary feedback data is positive in effect and the preliminary resource allocation strategy is close to the first resource allocation strategy, the original usage parameters are increased; on the contrary, if the preliminary resource allocation strategy is close to the second resource allocation strategy, the original usage parameters are decreased. On the other hand, if the preliminary feedback data is negative in effect and the preliminary resource allocation strategy is close to the first resource allocation strategy, the original usage parameters are decreased, while if the preliminary resource allocation strategy is close to the second resource allocation strategy, the original usage parameters are increased.
[0116] In step S402 of some embodiments, the preliminary resource allocation strategy is adjusted through the strategy adjustment system and the preliminary feedback data to obtain the third resource allocation strategy. It should be noted that the construction method of the third resource allocation strategy is the same as that of the first resource allocation strategy, which will not be elaborated here.
[0117] In step S403 of some embodiments, the preliminary resource allocation strategy is adjusted through the strategy adjustment model, the preliminary feedback data, and the target usage parameters to obtain the fourth resource allocation strategy. It should be noted that the construction method of the fourth resource allocation strategy is the same as that of the second resource allocation strategy, which will not be elaborated here.
[0118] In step S404 of some embodiments, the third resource allocation strategy and the fourth resource allocation strategy are fused according to the target usage parameters to obtain the target resource allocation strategy, and then the target resource allocation strategy replaces the preliminary resource allocation strategy as the current resource allocation strategy of the target resource allocation platform. It should be noted that the fusion method of the third resource allocation strategy and the fourth resource allocation strategy is the same as the fusion method of the first resource allocation strategy and the second resource allocation strategy above, which will not be elaborated here.
[0119] In steps S401 to S404 illustrated in this embodiment, after the joint resource adjustment system and the resource adjustment model generate new resource allocation strategies again with the preliminary resource allocation strategy and the preliminary feedback data, the two new resource allocation strategies are fused into a target resource allocation strategy and replace the preliminary resource allocation strategy to achieve the timed update of the resource allocation strategy, so as to continuously optimize a resource allocation strategy with better operation effect.
[0120] Please refer to Figure 5 , in some embodiments, the preliminary feedback data includes: platform usage growth data and platform resource loss data; step S401 may include but is not limited to steps S501 to S503:
[0121] Step S501, perform normalization processing on the platform usage growth data to obtain a platform usage growth parameter;
[0122] Step S502, perform normalization processing on the resource loss data to obtain a resource reduction parameter;
[0123] Step S503, adjust the original usage parameter according to the platform usage growth parameter and the resource reduction parameter to obtain a target usage parameter.
[0124] In step S501 of some embodiments, the platform usage growth data characterizes the growth of the access frequency of the target resource allocation platform. By means of web crawling, the access volume of the target resource allocation platform after executing the preliminary resource allocation strategy can be determined, and the access growth data can be calculated with the access volume under the original resource allocation strategy, so as to use the access growth data as the platform usage growth data. It should be noted that normalizing the platform usage growth data means changing the platform usage growth data into a value within the range of 0-1 as the platform usage growth parameter. For example, in the field of object management, the platform usage growth parameter can be the participation rate of logistics workers or the sign-in volume of logistics workers, and the participation rate or sign-in volume determined by obtaining the access volume of each logistics worker on the target resource allocation platform, or the number of logistics workers participating in the point redemption activity, so as to determine the participation rate growth rate according to the participation rate of logistics workers under the original resource allocation strategy and the preliminary resource allocation strategy. Then, the normalized participation rate growth rate of logistics workers is used as the reward coefficient under the preliminary resource allocation strategy.
[0125] In step S502 of some embodiments, the resource loss data is the amount of resource loss generated by the target resource allocation platform after using the preliminary resource allocation strategy. It should be noted that normalizing the resource loss data into a resource reduction parameter between 0-1 facilitates the adjustment of the original usage parameter. For example, in the field of object management, the resource loss data can be the cost increase rate. By normalizing the cost increase rate, a penalty coefficient is obtained, so as to determine whether the original usage parameter is lowered according to the penalty coefficient.
[0126] In step S503 of some embodiments, the platform usage growth parameter can characterize an increase in the original usage parameter, while the resource reduction parameter characterizes a decrease in the original usage parameter. Therefore, the adjustment direction and adjustment degree of the original usage parameter are determined according to the platform usage growth parameter and the resource reduction parameter to obtain the target usage parameter.
[0127] It should be noted that when adjusting the original usage parameters according to the platform usage growth parameter and the resource reduction parameter, it is also necessary to determine which training period the policy adjustment model belongs to. Because in the initial stage of training, the resource allocation policy mainly output by the policy adjustment coefficient is mainly used, and the resource allocation policy output by the policy adjustment model will only be added in the middle and late stages of training. Therefore, the target usage parameter and the preset threshold are used to determine which training period the policy adjustment model is in. In the initial stage of training, the usage parameter of the policy adjustment model is generated as shown in Equation (2):
[0128] β * =arg min β E (a,s)~A [L(π β (s),a)] (2)
[0129] β * is the usage parameter of the policy adjustment model. In the initial stage of training, β - is β 原 , (a, s)~A is the third resource allocation policy made based on the policy adjustment system, where s is the preliminary feedback data and a is the preliminary resource allocation policy. Therefore, first determine the platform usage growth parameter and the resource reduction parameter based on the preliminary feedback data, then determine the usage parameter of the policy adjustment model, and then determine the target usage parameter.
[0130] In the middle and late stages, the resource allocation policy output by the policy adjustment model and the resource allocation policy output by the policy adjustment system begin to approach. Therefore, the usage parameters of the policy adjustment model and the policy adjustment system can be determined by Equation (3):
[0131] α - ,β - =arg max α,β E s~B [R(F α,β (s),i,c))] (3)
[0132] In the formula, α - is the usage parameter of the policy adjustment system, β - =1-α - . In the initial stage of training, the usage parameter of the policy adjustment system is the original usage parameter, that is, α - =α 原 .
[0133] It should be noted that after determining the target usage parameter, the new target resource allocation policy can be determined, and the target resource allocation policy is determined as shown in Equation (4):
[0134] F(s)=α * π(s)+(1-α * )π β (s) (4)
[0135] In the middle and late stages of training, the α of the policy adjustment system * is α 目标 , and α 目标 is the target usage parameter, π(s) is the third resource allocation policy, and π β (s) is the fourth resource allocation policy.
[0136] For example, in the field of object management, the preliminary resource allocation policy is "at time t, the activity intensity is increased by 10%, and the commodity discount intensity is increased by 10%", and the preliminary feedback data generated by this preliminary resource allocation policy may be "the participation rate of the little brother is increased by 1%, and the actual cost is increased by 2%". Therefore, after the participation rate increase rate generated by the behavior group is standardized, it is the reward coefficient corresponding to the preliminary resource allocation policy, and the corresponding cost increase rate is the generated penalty coefficient.
[0137] In steps S501 to S503 shown in this embodiment, the platform usage growth parameter is determined through the platform usage growth data, the resource reduction parameter is determined through the platform resource loss data, and finally the original usage parameter is adjusted according to the platform usage growth parameter and the resource reduction parameter to obtain a more accurate target usage parameter.
[0138] Please refer to Figure 6 , in some embodiments, step S403 may include but is not limited to steps S601 to S603:
[0139] Step S601, obtaining the model loss data of the policy adjustment model according to the target usage parameter;
[0140] Step S602, adjusting the parameters of the policy adjustment model according to the model loss data;
[0141] Step S603, performing policy adjustment on the preliminary resource allocation policy and the preliminary feedback data through the adjusted policy adjustment model to obtain the fourth resource allocation policy.
[0142] In step S602 of some embodiments, because the reference data for adjusting the policy adjustment model is different in different periods, the model loss data of the policy adjustment model is determined according to the target usage parameter, and then it is determined which reference data the policy adjustment model is adjusted by. For example, in the initial stage of training, the policy adjustment model mainly focuses on imitating the policy adjustment coefficient, that is, making the resource allocation policy output by the policy adjustment model close to the resource allocation output by the policy adjustment coefficient to complete the initial training of the policy adjustment model. In the middle and late stages of training, the policy adjustment model already has professional policy adjustment experience. In this stage, it no longer mainly imitates the policy adjustment system, but is adjusted based on the collected feedback data to construct a resource allocation policy that more conforms to the operation effect of the target resource allocation platform.
[0143] In step S603 of some embodiments, the preliminary resource allocation policy is adjusted by the adjusted policy adjustment model and the preliminary feedback data to construct a better fourth resource allocation policy.
[0144] It should be noted that after the policy adjustment model outputs a resource allocation policy each time, steps S601 to S603 are repeatedly executed once to make the resource allocation policy output by the policy adjustment model more in line with the operation requirements of the target resource allocation platform.
[0145] In steps S601 to S603 illustrated in this embodiment, by determining which training stage the policy adjustment model is in, the corresponding model adjustment parameters are used to adjust the policy adjustment model, and then the policy is adjusted by the adjusted policy adjustment model to construct a resource allocation policy that better meets the operation requirements of the target resource allocation platform.
[0146] Please refer to Figure 7 , in some embodiments, step S601 includes but is not limited to steps S701 to S703:
[0147] Step S701, comparing the target usage parameter with a preset threshold to obtain a comparison result;
[0148] Step S702, if the comparison result indicates that the target usage parameter is greater than the preset threshold, the loss data between the first resource allocation policy and the second resource allocation policy is used as the model loss data;
[0149] Step S703, if the comparison result indicates that the target usage parameter is less than or equal to the preset threshold, the preliminary feedback data is used as the model loss data.
[0150] In step S701 of some embodiments, by comparing the target usage parameter with the preset threshold, it is determined which training stage the policy adjustment model is in. In this embodiment, if the target usage parameter is greater than the preset threshold, it is determined that the policy adjustment model is in the initial stage of training; if the target usage parameter is less than or equal to the preset threshold, it is determined that the policy adjustment model is in the middle and late stages of training. For example, in the initial stage of training, the value of the target usage parameter is between 0.7 and 1, and in the middle and late stages of training, the value of the target usage parameter is between 0.3 and 0.7. Therefore, it is easy to determine which training stage the policy adjustment model is in according to the target usage parameter and the preset threshold.
[0151] In step S702 of some embodiments, if the comparison result indicates that the target usage parameter is greater than the preset threshold, it is determined that the policy adjustment model is in the initial stage of training. Therefore, the policy adjustment model mainly imitates the policy adjustment system, and calculates the loss data between the first resource allocation policy and the second resource allocation policy as the model loss data. It should be noted that based on the loss data between the first resource allocation policy and the second resource allocation policy, the model parameters of the policy adjustment model are adjusted to make the resource allocation policy output by the policy adjustment model closer to the resource allocation policy output by the policy adjustment system.
[0152] In step S703 of some embodiments, if the comparison result indicates that the target usage parameter is less than or equal to the preset threshold, it indicates that the policy adjustment model is in the middle and late stages of training. It no longer mainly imitates the policy adjustment system, but adjusts the policy adjustment model according to the preliminary feedback data fed back by the target object to train a policy adjustment model that meets the operation requirements of the target resource allocation platform.
[0153] In steps S701 to S703 shown in this embodiment, by comparing the target usage parameter with the preset threshold, it is determined which training stage the policy adjustment model is in, and different data are set as model adjustment parameters to construct a policy adjustment model that more meets the operation requirements of the target resource allocation platform.
[0154] In some embodiments, the policy adjustment model includes at least one policy adjustment network, and the types of different policy adjustment networks are different. The policy adjustment network can be a linear regression network, a boosts tree network, a logistic regression network, etc. Please refer to Figure 8 , in this embodiment, it is set that the policy adjustment model is composed of three policy adjustment networks, namely a linear regression network, a boosts tree network, and a logistic regression network, and the linear regression network, the boosts tree network, and the logistic regression network are connected in parallel. Therefore, the training process of the policy adjustment model can be to recombine the policy adjustment network or to adjust the parameters of the policy adjustment network.
[0155] Please refer to Figure 9 , in some embodiments, step S603 may include but is not limited to steps S901 to S902:
[0156] Step S901, input the original resource allocation policy and the original feedback data into each adjusted policy adjustment network for policy adjustment to obtain at least one candidate resource allocation policy;
[0157] Step S902, perform a fusion process on at least one candidate resource allocation policy to obtain a fourth resource allocation policy.
[0158] In steps S901 to S902 of some embodiments, each policy adjustment network sets corresponding network weights. After each policy adjustment network outputs a candidate resource allocation policy, at least one candidate resource allocation policy is combined into a fourth resource allocation policy based on the network weights. It should be noted that in this embodiment, the fourth resource allocation policy is output after weighted summing at least one candidate resource allocation policy with network weights. In other embodiments, multiple candidate resource allocation policies can be directly summed or averaged to obtain the fourth resource allocation policy. The combination method for at least one candidate resource allocation policy is not specifically limited.
[0159] In steps S901 to S902 illustrated in this embodiment, the resource allocation policy is adjusted through different types of policy adjustment networks and preliminary feedback data to output candidate resource allocation policies, and then at least one candidate resource allocation policy is combined into a fourth resource allocation policy to obtain a better resource allocation policy.
[0160] As Figure 10 shown, in each embodiment of this application when adjusting the resource allocation policy, the feedback data of the target object based on the resource allocation policy is collected by the data monitoring module and defined as s, and then the feedback data s is input into the policy adjustment system to output the first resource allocation policy a1, and the feedback data is input into the policy adjustment model to output the second resource allocation policy a2. The first resource allocation policy a1 and the second resource allocation policy a2 are fused into the adjusted resource allocation policy A by using the parameter α. At the same time, after each resource allocation policy is completed, the reinforcement learning module adjusts the use of the parameter α according to the feedback data s, the first resource allocation policy a, and the second resource allocation policy a2. Specifically, the policy adjustment system inputs the feedback data s and the first resource allocation policy a1 into the reinforcement learning module, and the policy adjustment model outputs the feedback data and the second resource allocation policy and inputs them into the reinforcement learning module, so that the reinforcement learning module can adjust the use of the parameter α. Therefore, by jointly using the policy adjustment system with policy adjustment experience and the intelligent policy adjustment model, and relying on human experience in resource allocation policy adjustment to let the policy adjustment model overcome the cold start, and then quickly learn the experience in resource allocation policy adjustment during actual operation to adjust the parameters of the policy adjustment model, making the policy adjustment model the main force for policy adjustment in the middle and late stages, reducing the dependence on manpower, and realizing the joint completion of resource allocation policy adjustment by humans and machines, which can not only save manpower, improve the efficiency of resource allocation policy construction, but also reduce the time for resource allocation policy construction.
[0161] Please refer to Figure 11 , this application embodiment also provides a resource allocation policy adjustment device, which can implement the above resource allocation policy adjustment method. The device includes:
[0162] A policy acquisition module 1101, configured to acquire the original resource allocation policy of the target resource allocation platform;
[0163] An original feedback acquisition module 1102, configured to acquire the feedback data of the target object based on the original resource allocation policy to obtain the original feedback data; wherein, the original feedback data characterizes the feedback situation of the target object on the target resource allocation platform based on the original resource allocation policy;
[0164] A sending module 1103, configured to send the original resource allocation policy and the original feedback data to a preset policy adjustment system, and receive the first resource allocation policy fed back by the policy adjustment system according to the original resource allocation policy and the original feedback data;
[0165] A model adjustment module 1104, configured to perform policy adjustment on the original resource allocation policy according to a preset policy adjustment model and the original feedback data to obtain a second resource allocation policy;
[0166] A parameter acquisition module 1105, configured to acquire the original usage parameters output by the policy adjustment system, and perform fusion processing on the first resource allocation policy and the second resource allocation policy according to the original usage parameters to obtain a preliminary resource allocation policy; wherein, the original usage parameters characterize the usage probability of the first resource allocation policy;
[0167] A preliminary feedback acquisition module 1106, configured to acquire the feedback data of the target object based on the preliminary resource allocation policy to obtain the preliminary feedback data; wherein, the preliminary feedback data characterizes the feedback situation of the target object on the target resource allocation platform based on the preliminary resource allocation policy;
[0168] A policy adjustment module 1107, configured to adjust the preliminary resource allocation policy according to the preliminary feedback data.
[0169] The specific implementation manner of this resource allocation policy adjustment device is basically the same as the specific embodiments of the above resource allocation policy adjustment method, and will not be elaborated here.
[0170] An embodiment of this application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above resource allocation policy adjustment method is implemented. This electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0171] Please refer to Figure 12 , Figure 12 , which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0172] The processor 1201 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0173] The memory 1202 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1202 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1202 and are called by the processor 1201 to execute the resource allocation strategy adjustment method in the embodiments of the present application;
[0174] The input / output interface 1203 is used to implement information input and output;
[0175] The communication interface 1204 is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);
[0176] The bus 1205 transmits information between the various components of the device (such as the processor 1301, the memory 1202, the input / output interface 1203, and the communication interface 1204);
[0177] Among them, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 are communicatively connected to each other inside the device through the bus 1205.
[0178] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned resource allocation strategy adjustment method is implemented.
[0179] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0180] The resource allocation strategy adjustment method, device, computer device, and storage medium provided by the embodiments of the present application, by combining a strategy adjustment system with experience in strategy adjustment and an intelligent strategy adjustment model, rely on human experience in resource allocation strategy adjustment to enable the strategy adjustment model to overcome the cold start, and then quickly learn the experience in resource allocation strategy adjustment during actual operation to adjust the parameters of the strategy adjustment model, making the strategy adjustment model the main force in the mid- and late-stage strategy adjustment, reducing the dependence on manpower, and achieving the joint completion of resource allocation strategy adjustment by humans and machines. This not only saves manpower, improves the efficiency of resource allocation strategy construction, but also reduces the time for resource allocation strategy construction.
[0181] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0182] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0183] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0184] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0185] In the description of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0186] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0187] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0188] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0189] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0190] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0191] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A method for adjusting a resource allocation strategy, characterized in that The method includes: Obtaining the original resource allocation strategy of the target resource allocation platform; Obtaining the feedback data of the target object based on the original resource allocation strategy to obtain the original feedback data; wherein, the original feedback data represents the feedback of the target object on the target resource allocation platform based on the original resource allocation strategy; Sending the original resource allocation strategy and the original feedback data to a preset policy adjustment system, and receiving the first resource allocation strategy feedback by the policy adjustment system according to the original resource allocation strategy and the original feedback data; Performing policy adjustment on the original resource allocation strategy according to a preset policy adjustment model and the original feedback data to obtain a second resource allocation strategy; Obtaining the original usage parameters output by the policy adjustment system, and performing fusion processing on the first resource allocation strategy and the second resource allocation strategy according to the original usage parameters to obtain a preliminary resource allocation strategy; wherein, the original usage parameters represent the usage probability of the first resource allocation strategy; Obtaining the feedback data of the target object based on the preliminary resource allocation strategy to obtain the preliminary feedback data; wherein, the preliminary feedback data represents the feedback of the target object on the target resource allocation platform based on the preliminary resource allocation strategy; Adjusting the preliminary resource allocation strategy according to the preliminary feedback data.
2. The method according to claim 1, characterized in that, The adjusting the preliminary resource allocation strategy according to the preliminary feedback data includes: Performing adjustment processing on the original usage parameters according to the preliminary feedback data to obtain target usage parameters; Sending the preliminary resource allocation strategy and the preliminary feedback data to the policy adjustment system, and receiving the third resource allocation strategy feedback by the policy adjustment system according to the preliminary resource allocation strategy and the preliminary feedback data; Performing policy adjustment on the preliminary resource allocation strategy through the policy adjustment model, the target usage parameters and the preliminary feedback data to obtain a fourth resource allocation strategy; Adjusting the preliminary resource allocation strategy according to the target usage parameters, the third resource allocation strategy and the fourth resource allocation strategy.
3. The method according to claim 2, wherein The preliminary feedback data includes: platform usage growth data and platform resource loss data; the performing adjustment processing on the original usage parameters according to the preliminary feedback data to obtain target usage parameters includes: Performing standardization processing on the platform usage growth data to obtain platform usage growth parameters; Performing standardization processing on the resource loss data to obtain resource reduction parameters; Performing adjustment processing on the original usage parameters according to the platform usage growth parameters and the resource reduction parameters to obtain the target usage parameters.
4. The method according to claim 3, characterized in that, The performing policy adjustment on the preliminary resource allocation strategy through the policy adjustment model, the target usage parameters and the preliminary feedback data to obtain a fourth resource allocation strategy includes: According to the target usage parameters Obtaining the model loss data of the policy adjustment model; Performing parameter adjustment on the policy adjustment model according to the model loss data; Adjust the preliminary resource allocation strategy by using the adjusted policy adjustment model and the preliminary feedback data to obtain a fourth resource allocation strategy.
5. The method according to claim 4, characterized in that, Obtaining the model loss data of the policy adjustment model according to the target usage parameter includes: Compare the target usage parameter with a preset threshold to obtain a comparison result; If the comparison result indicates that the target usage parameter is greater than the preset threshold, use the loss data between the first resource allocation strategy and the second resource allocation strategy as the model loss data; If the comparison result indicates that the target usage parameter is less than or equal to the preset threshold, use the preliminary feedback data as the model loss data.
6. The method according to claim 4, characterized in that, The policy adjustment model includes at least one policy adjustment network; adjusting the preliminary resource allocation strategy by using the adjusted policy adjustment model and the preliminary feedback data to obtain a fourth resource allocation strategy includes: Input the original resource allocation strategy and the original feedback data into each adjusted policy adjustment network for policy adjustment to obtain at least one candidate resource allocation strategy; Perform a fusion process on at least one of the candidate resource allocation strategies to obtain the fourth resource allocation strategy.
7. The method according to any one of claims 1 to 6, characterized in that Obtaining the original usage parameter output by the policy adjustment system, and performing a fusion process on the first resource allocation strategy and the second resource allocation strategy according to the original usage parameter to obtain a preliminary resource allocation strategy includes: Obtain the original usage parameter output by the policy adjustment system; Extract the first resource allocation parameter from the first resource allocation strategy; Extract the second resource allocation parameter from the second resource allocation strategy; Select a target parameter adjustment operation from the preset candidate parameter adjustment operations according to the first resource allocation parameter and the second resource allocation parameter; Execute the target parameter adjustment operation on the first resource allocation parameter and the second resource allocation parameter according to the original usage parameter to obtain a target resource allocation parameter; Construct the preliminary resource allocation strategy according to the target resource allocation parameter.
8. A resource allocation strategy adjustment device, characterized in that, The device includes: A policy acquisition module for acquiring the original resource allocation strategy of the target resource allocation platform; An original feedback acquisition module for obtaining the feedback data of the target object based on the original resource allocation strategy to obtain the original feedback data; wherein, the original feedback data represents the feedback situation of the target object on the target resource allocation platform based on the original resource allocation strategy; A sending module for sending the original resource allocation strategy and the original feedback data to a preset policy adjustment system, and receiving the first resource allocation strategy fed back by the policy adjustment system according to the original resource allocation strategy and the original feedback data; A model adjustment module for adjusting the original resource allocation strategy according to a preset policy adjustment model and the original feedback data to obtain a second resource allocation strategy; A parameter acquisition module, configured to acquire the original usage parameters output by the policy adjustment system, and perform fusion processing on the first resource allocation policy and the second resource allocation policy according to the original usage parameters to obtain a preliminary resource allocation policy; wherein, the original usage parameters characterize the usage probability of the first resource allocation policy. A preliminary feedback acquisition module, configured to obtain feedback data of the target object based on the preliminary resource allocation policy to obtain preliminary feedback data; wherein, the preliminary feedback data characterizes the feedback of the target object on the target resource allocation platform based on the preliminary resource allocation policy. A policy adjustment module, configured to adjust the preliminary resource allocation policy according to the preliminary feedback data.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the resource allocation policy adjustment method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the resource allocation policy adjustment method according to any one of claims 1 to 7.
Citation Information
Cited By
Wireless communication resource allocation method and system based on artificial intelligence
CN121397753A