Business resource allocation method, system and device and storage medium
By collecting enterprise data and using reinforcement learning and game theory to generate resource allocation schemes, combined with manual review, the problem of low efficiency and lack of flexibility in resource allocation in existing technologies has been solved, achieving efficient and fair resource allocation decisions.
Patent Information
- Application Number
- CN202511212131.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-21
AI Technical Summary
Existing enterprise resource allocation methods are inefficient, lack objective and unified quantitative standards, are easily affected by human factors, cannot dynamically adapt to market changes, and the allocation logic of automated tools relies on fixed rules, resulting in insufficient flexibility.
By collecting enterprise operational status and historical data through data interfaces, resource requirements are generated using reinforcement learning models, preliminary plans are generated by combining game theory coordination mechanisms, and the final plans are determined through manual review. Resource allocation is carried out in a human-machine collaborative mode.
It has automated resource demand forecasting and allocation, improved planning efficiency and response speed, ensured the mathematical fairness of allocation schemes and optimal overall benefits, and enhanced the scientific nature of resource utilization and the quality of decision-making.
Smart Images

Figure CN120996498A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, specifically relating to a business resource allocation method, system, device, and storage medium. Background Technology
[0002] Current enterprise resource allocation, especially financial planning, largely relies on manual operations. Decision-makers typically make subjective judgments and manually allocate resources based on departmental applications, historical experience, and limited financial data. This approach has significant drawbacks: first, it is inefficient, struggling to quickly process complex applications from multiple departments and involving multiple factors; second, it lacks objective and unified quantitative standards, making it susceptible to human influence and lacking fairness and scientific rigor; and third, it cannot dynamically adapt to rapid changes in the market environment and internal operations, resulting in poor planning accuracy. Although some automated tools have been developed to assist in data aggregation, their core allocation logic still relies on preset, fixed rules, lacking flexibility and failing to achieve intelligent decision-making and multi-objective optimization. Therefore, how to efficiently, fairly, and intelligently complete resource allocation has become a pressing technical challenge in this field. Summary of the Invention
[0003] In view of the above-mentioned shortcomings of the prior art, the present invention provides a business resource allocation method, system, device and storage medium to solve the above-mentioned technical problems.
[0004] In a first aspect, the present invention provides a method for allocating business resources, comprising: The system collects operational status data and historical data from various business departments through data interfaces, and then performs data cleaning and formatting. Based on the aforementioned state data and historical data, a reinforcement learning model is used to generate business resource requirements for each business department. Based on the business resource needs of each business department, a preliminary resource allocation plan is generated using a game theory coordination mechanism. The preliminary resource allocation plan is sent to the manual review channel, and the resource allocation plan that has passed the manual review is determined as the final resource allocation plan.
[0005] In one optional implementation, status data of enterprise operations and historical data of various business departments are collected through a data interface, and the data is cleaned and formatted, including: Data is collected from various heterogeneous data sources, including databases, application programming interfaces, and file servers; the status data includes real-time cash positions and market environment data, and the historical data includes historical expenditures, revenues, and project urgency data for each department. The collected data is deduplicated, outliers are removed, or missing values are filled; the data is converted into a preset standardized format; and the data is aggregated or subjected to pivot calculations to generate specific dimensional indicators for model input.
[0006] In one optional implementation, the input to the reinforcement learning model includes: Historical data for each business unit, including historical expenditures, historical revenues, and project urgency indicators; Enterprise operational status data, including real-time cash position, market interest rates, and supply chain risk indicators; The output of the reinforcement learning model is the priority of each business department's funding needs and the suggested allocation amount; The reinforcement learning model is a value- or policy-based deep reinforcement learning network, whose optimization objective is to maximize long-term business benefits while satisfying the enterprise's total resource pool constraints.
[0007] In one optional implementation, the training method for the reinforcement learning model includes: Based on preset business rules and historical business data, the rule engine generates baseline resource requirement data for each business department. Using the historical business data and the corresponding baseline resource demand data as a training sample set, the initial reinforcement learning model is trained offline to obtain a primary model. The primary model is extrapolated onto historical data, and its output is compared and verified with the baseline resource demand data. Once the verification is successful, the initial model will be deployed in the production environment to generate business resource requirements. The business rules include at least one of the following: priority allocation rules based on project urgency, allocation rules based on the proportion of historical average expenditure, or allocation rules based on the order of business occurrence dates. The offline reinforcement learning algorithm used is one of conservative Q-learning, batch reinforcement learning, or model-based offline policy evaluation algorithm.
[0008] In one optional implementation, the model is deployed to the production environment via a phased rollout deployment; the phased rollout deployment includes: The model's output suggestions are mixed with the baseline results output by the rule engine in a certain proportion, and the mixing proportion of the model's suggestions is gradually increased.
[0009] In one optional implementation, a preliminary resource allocation plan is generated using a game theory coordination mechanism based on the business resource needs of each business department, including: Obtain the business resource requirements of multiple business departments, wherein the business resource requirements include at least the amount, priority indicators, historical contribution indicators, and urgency indicators; Based on a preset weight combination, the multiple application indicators of a single department or a departmental alliance are mapped to a total contribution value to define the characteristic function v(S) of the game. Based on the feature function v(S), the Monte Carlo sampling algorithm is used to approximate the Shapley value of each business unit in the resource allocation game. Based on the proportion of each business unit's Shapley value to the total Shapley value, resources in the total resource pool are allocated to generate a preliminary resource allocation plan.
[0010] In an optional implementation, the multiple application indicators of a single department or departmental alliance are mapped to a total contribution value according to a preset weight combination to define the characteristic function v(S) of the game, including:
[0011] Among them, the amount i Let w1 represent the quantified value of the monetary indicator for the i-th business unit, and w1 be the weight of the monetary indicator; priority i w2 represents the quantified value of the priority indicator for the i-th business unit, where w2 is the weight of the priority indicator; contribution. i w3 represents the quantified value of the contribution indicator for the i-th business unit, and w3 is the weight of the contribution indicator; urgency i w4 represents the urgency index of the i-th business unit, and w4 is the weight of the urgency index.
[0012] Secondly, the present invention provides a business resource allocation system, comprising: The data acquisition module is used to collect status data of enterprise operations and historical data of various business departments through data interfaces, and to perform data cleaning and formatting. The first allocation module is used to generate business resource requirements for each business department based on the state data and historical data, using a reinforcement learning model. The second allocation module is used to generate a preliminary resource allocation plan based on the business resource needs of each business department using a game theory coordination mechanism. The manual review module is used to send the preliminary resource allocation plan to the manual review channel and determine the resource allocation plan that has passed the manual review as the final resource allocation plan.
[0013] Thirdly, a device is provided, comprising: Memory, used to store business resource allocation programs; A processor is configured to implement the steps of the business resource allocation method as provided in the first aspect when executing the business resource allocation procedure.
[0014] Fourthly, a computer-readable storage medium is provided, on which a service resource allocation program is stored, and when the service resource allocation program is executed by a processor, it implements the steps of the service resource allocation method provided in the first aspect.
[0015] The beneficial effects of this invention are as follows: The business resource allocation method, system, equipment, and storage medium provided by this invention effectively overcome the subjectivity and inefficiency of traditional manual allocation methods by introducing an intelligent collaborative mechanism based on reinforcement learning and game theory. Its beneficial effects are: First, it automates resource demand prediction and allocation, significantly improving the efficiency and response speed of planning; second, it utilizes game theory mechanisms (such as Shapley values) to ensure the mathematical fairness and optimal overall benefit of the allocation scheme, overcoming human bias; finally, it innovatively adopts a human-machine collaborative model of "AI preliminary decision-making + human final review," which leverages the computational advantages of intelligent algorithms while retaining the supervision and control rights of human experts, ensuring the reliability and interpretability of the scheme, and ultimately significantly improving the scientific nature of enterprise resource utilization and the quality of decision-making. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic flowchart of a method according to an embodiment of the present invention.
[0018] Figure 2 This is a schematic block diagram of a system according to an embodiment of the present invention.
[0019] Figure 3 This is a schematic diagram of the structure of a device provided in an embodiment of the present invention. Detailed Implementation
[0020] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0022] The business resource allocation method provided in this embodiment of the invention is executed by a computer device, and correspondingly, the business resource allocation system runs on the computer device.
[0023] Figure 1 This is a schematic flowchart illustrating a method according to an embodiment of the present invention. Wherein, Figure 1 The executing entity can be a business resource allocation system. Depending on different needs, the order of steps in this flowchart can be changed, and some steps can be omitted.
[0024] like Figure 1 As shown, the method includes: S1. Collect status data of enterprise operations and historical data of various business departments through data interfaces, and perform data cleaning and formatting; S2. Based on the state data and historical data, use a reinforcement learning model to generate business resource requirements for each business department; S3. Based on the business resource needs of each business department, a preliminary resource allocation plan is generated using a game theory coordination mechanism; S4. The preliminary resource allocation plan is sent to the manual review channel, and the resource allocation plan that has passed the manual review is determined as the final resource allocation plan.
[0025] In one embodiment of the present invention, based on step S1, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.
[0026] S101. Collect data from multiple heterogeneous data sources, including databases, application programming interfaces, and file servers; the status data includes real-time cash positions and market environment data, and the historical data includes historical expenditures, revenues, and project urgency data of each department.
[0027] Real-time status data: Fund position data must include 12 core fields such as account balance (accurate to the cent), available credit, and frozen amount, transmitted in JSON format, and ensured time sequence consistency through timestamps (accurate to the millisecond); Market environment data must cover three major categories: macroeconomic indicators (GDP growth rate, CPI), industry data (competitive landscape, policy changes), and financial market data (interbank lending rates, exchange rate midpoint). Among them, quantitative indicators retain 6 significant digits, and qualitative descriptions (such as policy texts) are stored using UTF-8 encoding.
[0028] Historical data: Departmental income and expenditure data are collected in a five-tuple structure of "Department ID - Accounting Period - Income and Expenditure Type - Amount - Person in Charge". Historical expenditures need to be associated with the corresponding project code and budget item. Project urgency data is quantified using a "1-5 point system" (1 is the lowest urgency and 5 is the highest). Urgency assessment criteria (such as contract delivery deadline and customer level) are collected simultaneously, and metadata such as assessor and assessment time are recorded through data lineage tags.
[0029] S102. Perform deduplication, outlier removal, or missing value filling on the collected data; convert the data into a preset standardized format; perform aggregation or pivot calculation on the data to generate specific dimensional indicators for model input.
[0030] In the deduplication phase, a composite index of "business primary key + timestamp" is constructed, and duplicate records are removed using Spark distributed computing to ensure data uniqueness. Outlier handling combines the IQR rule (removing data outside Q1-1.5IQR to Q3+1.5IQR) with time-series STL decomposition to identify and mark outliers. Missing values are filled using KNN weighted imputation (for numerical values) and mode imputation (for categorical values), with linear interpolation supplemented for time-series data. In the standardization phase, time fields are standardized according to ISO 8601 format, and amounts are converted to Decimal (18,2) type. The aggregation layer aggregates data through three dimensions: time (day / week / month), department, and project, generating trend features (month-on-month growth rate), volatility features (standard deviation), and correlation features (Pearson coefficient). After VIF test (VIF<10) and Min-Max normalization, the data forms the model input.
[0031] In one embodiment of the present invention, based on step S2, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.
[0032] First, the state space of the reinforcement learning model is constructed. This state space is a multi-dimensional continuous vector representing the system environment at time t. This state vector s... t It consists of the following normalized or standardized dimensions: the historical expenditure ratio of each department (e.g., the average of the past three months); the historical revenue generation ratio of each department; the urgency score of each department's current projects (mapped to the [0,1] range); the company's real-time cash position (the ratio of available funds to the total cash pool); market interest rates (normalized to Z-score); supply chain risk indicators (probability values between 0 and 1); and time period information (e.g., the normalized value for the current quarter).
[0033] Secondly, the model's action space is defined. These actions are used to generate resource requirement suggestions for each department, employing a hybrid action space combining discrete and continuous methods. For each department, the model outputs two actions: Priority action: A discrete value, selected from 5 predefined priority levels (e.g., levels 1-5, corresponding to normalized values of 0.2, 0.4, 0.6, 0.8, 1.0).
[0034] Allocation quota action: A continuous value representing the suggested resource quota to be allocated to this department. The Softmax function processes the proportional actions for all departments to meet the constraints of the overall resource pool.
[0035] Next, a reward function is designed to guide the model in learning the optimal policy. The reward r_t is a weighted sum of multiple objectives, calculated as follows: r t = W1·R 收益 +W2·R 风险 +W3·R 公平 +W4·R 约束 in: R 收益 The revenue-generating incentives are proportional to each department's revenue-generating capacity and the proportion allocated to them. R 风险 Risk penalties are proportional to the proportion allocated to supply chain risks and high-risk sectors. R 公平 To ensure fair rewards, the Gini coefficient is used to measure the fairness of the distribution ratio; R 约束 To constrain penalties, a large negative reward is applied when the total allocation ratio exceeds 1; W1, W2, W3, and W4 are preset weighting coefficients used to balance the importance of different objectives.
[0036] Then, a deep neural network is constructed as a function approximator for the reinforcement learning model. Value-based algorithms (such as DQN) or policy-based algorithms (such as PPO) can be selected based on specific needs. If DQN is used, the network architecture can contain a common hidden layer, which then branches into two output heads: one output layer generates Q-values for priority actions for each department (discrete); the other output layer generates suggested allocation quotas through Softmax (continuous).
[0037] If PPO is used, the policy network can output a Gaussian distribution (for generating continuous proportional actions) and logits of priority actions (for discrete classification).
[0038] The phased training process of reinforcement learning: First, a training baseline is established. Since training a reinforcement learning model directly from scratch carries the risk of generating unreasonable requirements, this invention first utilizes pre-defined business rules and historical business data to generate baseline resource requirement data for each business department through a rule engine. The rule engine incorporates various configurable business logics, such as: Prioritization rule based on project urgency: Departments with projects marked as high urgency will be allocated higher resource priority and proportion; Resource allocation is based on the historical average expenditure ratio: resources are allocated proportionally to each department based on the ratio of its average expenditure over a past period to its total expenditure. Allocation rules based on the order of business occurrence date: For payment-related requests, the earlier the business occurrence date, the higher the allocation priority.
[0039] The rules engine processes historical business data based on a combination of one or more of the aforementioned rules, outputting a set of reliable baseline requirements data that conforms to the company's basic operational logic. This data serves as the "gold standard" and safety barrier for model learning.
[0040] Secondly, offline reinforcement learning (Offline RL) training is performed. The historical business data (as state sequences) and the corresponding baseline resource requirement data (as action labels) together constitute an offline training sample set. Using an offline reinforcement learning algorithm, the initial reinforcement learning model is trained using this sample set without online interaction with the real business environment. This method avoids the direct impact of a poorly performing model in the early stages of training on the production system. The preferred offline reinforcement learning algorithm is Conservative Q-Learning (CQL), which adds a conservative regularization term to the learning objective of the Q-function to suppress overestimation of potentially high-risk actions that have not appeared in historical data, thereby training a more cautious and safer initial model. Batch reinforcement learning (BRL) or model-based offline policy evaluation (MBOPE) algorithms can also be used to achieve this goal.
[0041] Next, model validation and security testing are conducted. The trained preliminary model is run on another set of historical data (simulated operation), and its output resource requirement suggestions are compared with the baseline requirement data generated by the rule engine for validation. Validation metrics include mean absolute error (MAE) and root mean square error (RMSE). An error threshold is set; only when the error between the model's output and the baseline data is below this threshold is the model considered to have initially learned the effective strategies inherent in the rule engine, and the validation is considered successful.
[0042] Finally, implement a phased deployment. Once the model is validated, deploy it to the production environment. Initial deployment can use Shadow Mode, where the model runs in parallel with the existing rules engine. Its output demand suggestions are used only for monitoring and logging, not for actual decision-making, allowing for further observation of its performance in a completely real data stream. After a period of stable operation, gradually switch to a model-generated business resource requirements, ultimately achieving intelligent upgrades.
[0043] In one embodiment of the present invention, based on step S3, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.
[0044] S301. Obtain the business resource requirements of multiple business departments, wherein the business resource requirements include at least the amount, priority indicators, historical contribution indicators, and urgency indicators.
[0045] To address scenarios involving multi-departmental resource competition, a four-dimensional quantitative indicator system is established to achieve standardized representation of demand. Amount Indicators (Amount) i ): The amount of resources applied for by the i-th department (unit: RMB 10,000), based on the business resource demand extracted from the output of the reinforcement learning model.
[0046] Priority indicators (priority) i ): The priority of the reinforcement learning model output is quantified using a 1-10 score system (10 points is the highest priority).
[0047] Historical contribution index (contribution) i : Defined as the ratio of the department's cumulative net profit over the past three years to the company's total net profit. The value ranges from [1, 30] and is calculated using financial annual report data and processed by a moving average (window size = 4 quarters).
[0048] Urgency index (urgency) i The decay function is calculated based on the time difference between the business deadline and the current time. The formula is as follows: d represents the remaining number of days, with a value range of [1,5]. When d ≤ 7 days, the value is 5, and when d ≥ 90 days, the value is 1.
[0049] S302. Based on a preset weight combination, map the multiple application indicators of a single department or departmental alliance to a total contribution value to define the characteristic function v(S) of the game.
[0050] The four-dimensional indicators are uniformly mapped to the [0,1] interval through normalization.
[0051] A combined weighting strategy using the Analytic Hierarchy Process (AHP) and entropy weighting method is employed. Subjective weights: Calculated using a pairwise comparison matrix of enterprise management (priority w2=0.35, urgency w4=0.30, contribution w3=0.20, amount w1=0.15). Objective weighting: based on indicator information entropy ( The calculation shows that the final combined weights satisfy... .
[0052]
[0053] Among them, the amount i Let w1 represent the quantified value of the monetary indicator for the i-th business unit, and w1 be the weight of the monetary indicator; priority i w2 represents the quantified value of the priority indicator for the i-th business unit, where w2 is the weight of the priority indicator; contribution. i w3 represents the quantified value of the contribution indicator for the i-th business unit, and w3 is the weight of the contribution indicator; urgency i w4 represents the urgency index of the i-th business unit, and w4 is the weight of the urgency index.
[0054] S303. Based on the feature function v(S), the Monte Carlo sampling algorithm is used to approximate the Shapley value of each business unit in the resource allocation game; Formula for calculating Shapley value:
[0055] Where N represents the set of all departments, and S represents any subset that does not contain department i. Let n represent the number of departments in alliance S, and n represent the number of departments in set N. This represents the "characteristic function value" of the consortium S. This indicates the contribution level of the new alliance after joining department i. Weighting factors are used to ensure that all possible alliance orders are considered fairly.
[0056] Taking the calculation of the Shapley value for the sales department as an example: List all possible alliance combinations (N = {Sales, Purchasing, R&D, Administration}): {Procurement}, {Research and Development}, {Administration}, {Procurement, Research and Development}, {Procurement, Administration}, {Research and Development, Administration}, {Procurement, Research and Development, Administration}; For S = {procurement, R&D}, calculate the marginal contribution: v(S∪{Sales}) - v(S) = v({Sales,Purchasing,R&D}) - v({Purchasing,R&D}).
[0057] Shapley value for the sales department: ϕ 销售 = (1 / 4)*v({sales}) + (1 / 12)*(v({sales, purchases})-v({purchases}) + ... ).
[0058] S304. Based on the proportion of each business department's Shapley value to the total Shapley value, allocate resources in the total resource pool to generate a preliminary resource allocation plan.
[0059] Allocation ratio calculation: The resource allocation ratio of the i-th department is
[0060] satisfy This ensures the uniformity of the allocation.
[0061] Resource quantity determination: Let the total resource pool of the enterprise be C (unit: 10,000 yuan), then the allocation quantity for the i-th department is... At the same time, it is necessary to meet the department's application amount constraints ( If it exceeds the limit, then press Adjust and reallocate remaining resources (by r) j (The proportion is allocated to other departments).
[0062] Output: Generate a structured solution containing "Department ID - Allocation Amount - Percentage - Shapley Value - Constraint Satisfaction Status", along with fairness and efficiency indicators, providing quantitative basis for manual review.
[0063] In one embodiment of the present invention, based on step S4, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.
[0064] S401. Build a web-based manual review platform (based on Spring Boot backend + Vue.js frontend architecture) to achieve visual display and interactive review of the initial solution. Core functions include: Solution visualization: ECharts charts are used to display the global distribution of resource allocation (pie charts show the resource share of each department), the supply and demand matching of various resources (bar charts compare the allocated amount with the total stock), and the departmental revenue forecast (line charts show the trend of departmental revenue after the implementation of the solution). Drill-down is also supported to view detailed data (such as the difference between the historical allocation record of a certain type of resource in a department and the current allocation). Review process control: Based on the Activiti workflow engine, a multi-level review process is designed, with three review nodes: "preliminary review by department head - secondary review by resource management department - final review by enterprise decision-making level". The review time limit for each node is set to 24 hours. If the review is not completed within the time limit, an automatic reminder will be triggered (via email and SMS). Solution adjustment interface: Allows reviewers to adjust the allocation plan within their authorized scope (such as modifying the resource allocation of a department). The system verifies in real time whether the adjusted plan meets the constraints (if a resource conflict occurs after the adjustment, an early warning will pop up immediately and the reason for the conflict will be displayed).
[0065] S402. Finalization and Implementation of the Solution Review comments integration: The platform automatically summarizes the review comments from each stage and forms a structured report consisting of "review comments - adjustment suggestions - adjustment basis"; Final solution determination: If the review is approved (all nodes agree and there are no adjustment suggestions), the preliminary solution will be directly designated as the final solution; if there are adjustment suggestions and the adjusted solution meets the constraints, a revised final solution will be generated; if the review is not approved (there are major conflicts or strategic inconsistencies), the process will return to the S2 module to regenerate resource requirements. Solution Output and Archiving: The final solution is output to each business department in PDF format (including electronic signature) and simultaneously archived in the Enterprise Document Management System (EDMS). At the same time, allocation instructions are sent to the resource management system (such as personnel allocation in the HR system and fund transfer in the ERP system) through the API interface to realize the automated implementation of the solution.
[0066] In some embodiments, the business resource allocation system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the business resource allocation system may be stored in the memory of a computer device and executed by at least one processor to perform (see details). Figure 1 (Description) Functionality of business resource allocation.
[0067] In this embodiment, the business resource allocation system can be divided into multiple functional modules based on the functions it performs, such as... Figure 2As shown. The module referred to in this invention is a series of computer program segments that can be executed by at least one processor and perform a fixed function, and is stored in memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0068] The data acquisition module is used to collect status data of enterprise operations and historical data of various business departments through data interfaces, and to perform data cleaning and formatting. The first allocation module is used to generate business resource requirements for each business department based on the state data and historical data, using a reinforcement learning model. The second allocation module is used to generate a preliminary resource allocation plan based on the business resource needs of each business department using a game theory coordination mechanism. The manual review module is used to send the preliminary resource allocation plan to the manual review channel and determine the resource allocation plan that has passed the manual review as the final resource allocation plan.
[0069] Figure 3 The business resource allocation method provided in the embodiments of this application can be applied to devices. Those skilled in the art will understand that the device structure involved in the embodiments of this invention does not constitute a limitation on the device. A device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. In the embodiments of this invention, the device includes, but is not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of this application described and / or claimed herein.
[0070] The device 300 may include a processor 310, a memory 320, and a communication unit 330. These components communicate via one or more buses. Those skilled in the art will understand that the server structure shown in the figure does not constitute a limitation of the present invention. It may be a bus topology or a star topology, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0071] The memory 320 can be used to store execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the device 300 is able to perform some or all of the steps in the above method embodiments.
[0072] The processor 310 serves as the control center of the storage device, connecting various parts of the electronic device via various interfaces and lines. It executes software programs and / or modules stored in the memory 320, and calls data stored in the memory to perform various functions of the electronic device and / or process data. The processor can be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions connected together. For example, the processor 310 may consist only of a central processing unit (CPU). In this embodiment of the invention, the CPU may have a single processing core or include multiple processing cores.
[0073] The communication unit 330 is used to establish a communication channel, enabling the storage device to communicate with other devices. It can receive user data sent by other devices or send user data to other devices.
[0074] The present invention also provides a computer storage medium, wherein the computer storage medium may store a program, which, when executed, may include some or all of the steps provided in the embodiments of the present invention. The storage medium may be a magnetic disk, an optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0075] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other medium capable of storing program code. It includes several instructions to cause a computer device (which may be a personal computer, a server, or a second device, network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0076] The same or similar parts between the various embodiments in this specification can be referred to mutually. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.
[0077] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or modules may be electrical, mechanical, or other forms.
[0078] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0079] In addition, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0080] Although the present invention has been described in detail with reference to the accompanying drawings and preferred embodiments, the present invention is not limited thereto. Various equivalent modifications or substitutions can be made to the embodiments of the present invention by those skilled in the art without departing from the spirit and essence of the invention, and such modifications or substitutions should all be within the scope of the present invention. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should also be covered within the protection scope of the present invention.
Claims
1. A method for allocating business resources, characterized in that, include: The system collects operational status data and historical data from various business departments through data interfaces, and then performs data cleaning and formatting. Based on the aforementioned state data and historical data, a reinforcement learning model is used to generate business resource requirements for each business department. Based on the business resource needs of each business department, a preliminary resource allocation plan is generated using a game theory coordination mechanism. The preliminary resource allocation plan is sent to the manual review channel, and the resource allocation plan that has passed the manual review is determined as the final resource allocation plan.
2. The method according to claim 1, characterized in that, The system collects operational status data and historical data from various business departments through data interfaces, and performs data cleaning and formatting, including: Data is collected from various heterogeneous data sources, including databases, application programming interfaces, and file servers; the status data includes real-time cash positions and market environment data, and the historical data includes historical expenditures, revenues, and project urgency data for each department. The collected data is deduplicated, outliers are removed, or missing values are filled; the data is converted into a preset standardized format; and the data is aggregated or subjected to pivot calculations to generate specific dimensional indicators for model input.
3. The method according to claim 1, characterized in that, The inputs to the reinforcement learning model include: Historical data for each business unit, including historical expenditures, historical revenues, and project urgency indicators; Enterprise operational status data, including real-time cash position, market interest rates, and supply chain risk indicators; The output of the reinforcement learning model is the priority of each business department's funding needs and the suggested allocation amount; The reinforcement learning model is a value- or policy-based deep reinforcement learning network, whose optimization objective is to maximize long-term business benefits while satisfying the enterprise's total resource pool constraints.
4. The method according to claim 3, characterized in that, Training methods for reinforcement learning models include: Based on preset business rules and historical business data, the rule engine generates baseline resource requirement data for each business department. Using the historical business data and the corresponding baseline resource demand data as a training sample set, the initial reinforcement learning model is trained offline to obtain a primary model. The primary model is extrapolated onto historical data, and its output is compared and verified with the baseline resource demand data. Once the verification is successful, the initial model will be deployed in the production environment to generate business resource requirements. The business rules include at least one of the following: priority allocation rules based on project urgency, allocation rules based on the proportion of historical average expenditure, or allocation rules based on the order of business occurrence dates. The offline reinforcement learning algorithm used is one of conservative Q-learning, batch reinforcement learning, or model-based offline policy evaluation algorithm.
5. The method according to claim 4, characterized in that, The model is deployed to the production environment via a phased rollout deployment; the phased rollout deployment includes: The model's output suggestions are mixed with the baseline results output by the rule engine in a certain proportion, and the mixing proportion of the model's suggestions is gradually increased.
6. The method according to claim 1, characterized in that, Based on the business resource needs of each business department, a preliminary resource allocation plan is generated using a game theory coordination mechanism, including: Obtain the business resource requirements of multiple business departments, wherein the business resource requirements include at least the amount, priority indicators, historical contribution indicators, and urgency indicators; Based on a preset weight combination, the multiple application indicators of a single department or a departmental alliance are mapped to a total contribution value to define the characteristic function v(S) of the game. Based on the feature function v(S), the Monte Carlo sampling algorithm is used to approximate the Shapley value of each business unit in the resource allocation game. Based on the proportion of each business unit's Shapley value to the total Shapley value, resources in the total resource pool are allocated to generate a preliminary resource allocation plan.
7. The method according to claim 6, characterized in that, Based on a preset weight combination, the multiple application indicators of a single department or departmental alliance are mapped to a total contribution value to define the characteristic function v(S) of the game, including: Among them, the amount i Let w1 represent the quantified value of the monetary indicator for the i-th business unit, and w1 be the weight of the monetary indicator; priority i w2 represents the quantified value of the priority indicator for the i-th business unit, where w2 is the weight of the priority indicator; contribution. i w3 represents the quantified value of the contribution indicator for the i-th business unit, and w3 is the weight of the contribution indicator; urgency i w4 represents the urgency index of the i-th business unit, and w4 is the weight of the urgency index.
8. A business resource allocation system, characterized in that, include: The data acquisition module is used to collect status data of enterprise operations and historical data of various business departments through data interfaces, and to perform data cleaning and formatting. The first allocation module is used to generate business resource requirements for each business department based on the state data and historical data, using a reinforcement learning model. The second allocation module is used to generate a preliminary resource allocation plan based on the business resource needs of each business department using a game theory coordination mechanism. The manual review module is used to send the preliminary resource allocation plan to the manual review channel and determine the resource allocation plan that has passed the manual review as the final resource allocation plan.
9. A business resource allocation device, characterized in that, include: Memory, used to store business resource allocation programs; A processor, configured to implement the steps of the service resource allocation method as described in any one of claims 1-7 when executing the service resource allocation procedure.
10. A computer-readable storage medium storing a computer program, characterized in that, The readable storage medium stores a service resource allocation program, which, when executed by a processor, implements the steps of the service resource allocation method as described in any one of claims 1-7.