A resource allocation method and system in a human-machine collaborative decision-making scenario

Through a dynamic scoring mechanism based on task complexity, knowledge density and intelligent adaptability, the problem of low task matching efficiency in human-machine collaborative decision-making is solved, the rational allocation of resources and the improvement of task execution efficiency are achieved, and the needs of different task scenarios are adapted.

CN120235427BActive Publication Date: 2025-09-12CHINA SHENHUA ENERGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510715298.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In human-machine collaborative decision-making scenarios, existing technologies have low task matching efficiency, are unable to systematically integrate knowledge resources, and have unclear human-machine role positioning, resulting in poor performance and insufficient adaptation across cross-task scenarios.

Method used

Based on the three dimensional indicators of task complexity, knowledge density and intelligent adaptability, the weight coefficient is dynamically optimized through the preset strategy network and reinforcement learning model to generate a comprehensive task score, and the task is dynamically assigned to artificial or intelligent systems according to the priority of the score.

Benefits of technology

It improves the dynamic adaptability, decision-making reliability and task execution efficiency of the human-machine collaborative system, rationally allocates resources, avoids waste or excessive use, and enhances the versatility across task scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235427B_ABST
    Figure CN120235427B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of human-computer collaborative interaction, and discloses a resource allocation method and system in a human-computer collaborative decision-making scenario, the method comprising: generating corresponding task characteristics based on three dimensional indicators of task complexity, knowledge density, and intelligent adaptability of the task to be processed; initializing initial weight coefficients of the three dimensional indicators, and dynamically optimizing the weight coefficients corresponding to the three dimensional indicators according to the task characteristics and real-time system resource status using a preset strategy network; weighted normalizing the three dimensional indicators according to the weight coefficients to generate a comprehensive task score; dividing task priorities according to the comprehensive task score, and dynamically allocating tasks to manual and / or intelligent systems based on the priorities, which can give full play to the flexibility of manual experience and the efficiency of intelligent systems, flexibly adjust the division of labor between man and machine according to task characteristics, improve the versatility across task scenarios, and enhance the dynamic adaptability, decision reliability, and task execution efficiency of the human-computer collaborative system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer collaborative interaction, and in particular to a resource allocation method and system in a human-computer collaborative decision-making scenario. Background Art

[0002] In modern, complex human-machine collaborative decision-making scenarios, such as report compilation and fault diagnosis, traditional task allocation mechanisms rely heavily on human experience and judgment. This model has significant flaws: Due to the lack of quantitative evaluation criteria during task matching, the current system primarily relies on a rule base based on expert experience to allocate tasks. This makes it difficult to quickly and accurately match tasks with executors, resulting in inefficient task matching. Furthermore, it fails to systematically integrate and utilize internal and external knowledge resources, hindering knowledge reuse and circulation, leading to a serious lack of knowledge utilization.

[0003] Furthermore, during task execution, the roles of human and intelligent systems are often blurred, leading to extreme situations where human intervention is overly dependent or intelligent rules are mechanically enforced, resulting in an imbalance between human and machine roles. Existing human-machine collaborative decision-making methods suffer from static scoring, resulting in poor performance and insufficient adaptability to single scenarios. Therefore, a solution is needed that can fully leverage the flexibility of human experience, the efficiency of intelligent systems, and their versatility across task scenarios to improve the overall effectiveness of collaborative decision-making. Summary of the Invention

[0004] In view of this, the present invention provides a resource allocation method and system in a human-machine collaborative decision-making scenario, which solves the problem of how to give full play to the flexibility of human experience and the efficiency of intelligent systems and improve the versatility across task scenarios, and can significantly improve the dynamic adaptability, decision-making reliability and task execution efficiency of the human-machine collaborative system.

[0005] In a first aspect, the present invention provides a resource allocation method in a human-machine collaborative decision-making scenario, comprising:

[0006] Generate corresponding task features based on the three dimensional indicators of task complexity, knowledge density, and intelligent adaptability of the task to be processed;

[0007] Initialize the initial weight coefficients of the three-dimensional indicators, and use the preset strategy network to dynamically optimize the weight coefficients corresponding to the three-dimensional indicators according to task characteristics and real-time system resource status, including:

[0008] The task characteristics and system resource status corresponding to the three dimensional indicators of the task to be processed are used as the state space;

[0009] The weight coefficients α, β, and γ corresponding to the three dimensional indicators are used as the action space;

[0010] Using a preset reinforcement learning model, the weight coefficients corresponding to the three dimensions are dynamically optimized with a preset reward function, where the preset reward function η = accuracy × efficiency - cost, and the weight coefficients α, β, and γ satisfy the constraint α + β + γ = 1;

[0011] The three-dimensional indicators are weighted and normalized according to the weight coefficients output by the preset strategy network to generate a comprehensive task score;

[0012] Tasks are prioritized based on their comprehensive scores and dynamically assigned to manual and / or intelligent systems for execution based on priority.

[0013] The embodiment of the present invention generates task features through three dimensional indicators: task complexity, knowledge density, and intelligent adaptability. It can comprehensively and meticulously describe the essential characteristics of the task. The preset strategy network can learn the optimal weight combination under different task characteristics and system resource states, thereby making task allocation more reasonable. The three dimensional indicators are weighted and normalized to generate a comprehensive task score, providing a unified evaluation standard for all tasks. Regardless of the type and characteristics of the task, its priority in task allocation can be measured by a comprehensive score, making different tasks comparable. Task priorities are divided according to the comprehensive task score, and dynamically allocated to manual and / or intelligent systems to perform tasks based on the priority. This can give full play to the respective advantages of manual and intelligent systems, reasonably allocate resources, avoid waste or excessive use of resources, improve the versatility across task scenarios, and significantly improve the dynamic adaptability, decision reliability and task execution efficiency of the human-computer collaborative system.

[0014] In an optional embodiment, the task complexity is calculated by weighting three indicators: the number of task steps, the density of decision points, and the information dependency;

[0015] The knowledge density is calculated based on the weighted calculation of three indicators: knowledge base coverage, professional knowledge depth, and knowledge relevance;

[0016] The intelligent adaptability is calculated based on the weighted calculation of three indicators: process automation rate, algorithm accuracy, and human-computer interaction friendliness.

[0017] This embodiment of the present invention breaks down each dimension into multiple, detailed indicators and applies them in a weighted calculation, avoiding the one-sidedness of single-indicator evaluation and making task evaluation more realistic. Different indicators reflect task characteristics from different perspectives, and by comprehensively considering them in a weighted manner, a more comprehensive and objective evaluation system can be constructed.

[0018] In an optional embodiment, the triggering condition for adjusting the preset weight coefficient utilizes a preset strategy network based on task characteristics and real-time system resource status, and the triggering condition for adjusting the weight coefficient includes at least one of the following:

[0019] Time-driven, triggering weight adjustments at preset time intervals;

[0020] Performance-driven, weight adjustment is triggered when at least one of the accuracy, efficiency or cost indicators exceeds a preset threshold;

[0021] Task-driven, when the feature distribution of a new task deviates from historical data, it triggers weight adjustment;

[0022] Event-driven, weight adjustment is triggered when the system resource status changes or business priorities change.

[0023] The embodiment of the present invention prevents the weight coefficient from gradually deviating from the optimal value due to long-term operation of the system through regular adjustments, avoids large fluctuations in system performance caused by small changes accumulated over a long period of time, and helps maintain the stability and reliability of system performance; when a key performance indicator is abnormal, the system can quickly adjust the weight coefficient and try to find a new balance to solve the performance problem and prevent the problem from further deteriorating and causing a greater impact on the system; when the characteristic distribution of a new task deviates from historical data, the weight adjustment is triggered, allowing the system to quickly adapt to new types of tasks; when the system resource status changes, the weight adjustment is triggered, allowing the system to adjust the task processing strategy in time according to the availability of resources; changes in business priorities may cause changes in the importance of tasks. Through event-driven methods, when the business priority changes, the weight adjustment is triggered, and the system can quickly shift resources and attention to more important tasks, ensuring that the key goals of the business are achieved first, and improving the system's response speed and support capabilities for business needs.

[0024] In an optional embodiment, the prioritizing tasks according to their comprehensive scores and dynamically allocating tasks to manual and / or intelligent systems based on their priorities includes:

[0025] Divide the task's comprehensive score into three threshold intervals based on the business needs or historical data of the task to be processed. Different threshold intervals correspond to different priorities. Based on the mapping relationship between the preset priority and task allocation and the threshold interval in which the task's comprehensive score falls, determine whether to execute the task manually, using an intelligent system, or using human-machine collaboration; or

[0026] When there are multiple tasks to be processed, the comprehensive scores of the tasks to be processed are compared, and the task allocation method is determined based on the comparison results.

[0027] The embodiment of the present invention determines the threshold interval based on the business needs or historical data of the task to be processed, and can closely fit the actual business scenario. Different business scenarios have different requirements for tasks. Doing so can ensure that the division of task priorities meets the specific needs of the business, so that resources can be allocated first to tasks that have the greatest impact on the business, improving the overall operational efficiency and quality of the business. The mapping relationship between preset priorities and task allocations clarifies the task execution methods that should be adopted under different priorities and standardizes the task allocation process; by comparing their task comprehensive score values ​​to determine the task allocation method, tasks can be flexibly sorted according to specific circumstances. This method can dynamically adapt to changes and diversity in tasks, avoiding the irrationality that may be caused by fixed priority settings.

[0028] In an optional embodiment, the method further includes:

[0029] Generate task features based on the time sensitivity and resource load indicators of the tasks to be processed;

[0030] The initial weight coefficients of the five dimensions of task complexity, knowledge density, intelligent adaptability, time sensitivity and resource load of the task to be processed are initialized using the hierarchical analysis method or entropy weight method, and a real-time feedback loop is constructed to dynamically optimize the weight coefficients through the preset strategy network.

[0031] The embodiment of the present invention combines time sensitivity and resource load with task complexity, knowledge density, and intelligent adaptability, comprehensively covering task attributes and execution environment factors, helping to grasp the overall picture of the task more accurately, providing a more reliable basis for resource allocation, and ensuring that the resource allocation strategy can simultaneously meet task characteristics and system resource conditions; using the hierarchical analysis method or the entropy weight method to initialize the weight coefficient, providing an objective and reasonable starting point for task allocation, constructing a real-time feedback loop and combining it with a preset strategy network, so that the weight coefficient can be dynamically adjusted with task execution and environmental changes, so that the system can quickly adapt to task requirements and environmental changes.

[0032] In a second aspect, the present invention provides a resource allocation system in a human-machine collaborative decision-making scenario, comprising:

[0033] A three-dimensional evaluation model building module is used to generate corresponding task features based on the three dimensional indicators of task complexity, knowledge density, and intelligent adaptability of the task to be processed;

[0034] The dynamic weight adjustment module is used to initialize the initial weight coefficients of the three-dimensional indicators and dynamically optimize the weight coefficients corresponding to the three-dimensional indicators based on the preset strategy network according to the task characteristics and real-time system resource status;

[0035] The task comprehensive score calculation module is used to weight and normalize the three-dimensional indicators according to the weight coefficients output by the preset strategy network to generate a comprehensive task score;

[0036] The task allocation module is used to prioritize tasks based on their comprehensive scores and dynamically allocate tasks to manual and / or intelligent systems based on their priorities.

[0037] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions being stored in the memory, and the processor executing the computer instructions to thereby execute the resource allocation method in the human-computer collaborative decision-making scenario of the above-mentioned first aspect or any corresponding embodiment thereof.

[0038] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the resource allocation method in a human-computer collaborative decision-making scenario of the above-mentioned first aspect or any corresponding embodiment thereof.

[0039] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the resource allocation method in a human-computer collaborative decision-making scenario according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0041] Figure 1 is a flow chart of a resource allocation method in a human-machine collaborative decision-making scenario according to an embodiment of the present invention;

[0042] Figure 2 is a flow chart of a resource allocation method in another human-machine collaborative decision-making scenario according to an embodiment of the present invention;

[0043] Figure 3 This is a structural block diagram of a resource allocation system in a human-machine collaborative decision-making scenario according to an embodiment of the present invention;

[0044] Figure 4 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0046] In order to overcome the shortcomings of the existing technology that fail to give full play to the respective advantages of artificial and intelligent systems, resulting in limited task processing efficiency and quality, the embodiment of the present invention provides a resource allocation method in a human-machine collaborative decision-making scenario. Figure 1 is a flow chart of a resource allocation method in a human-machine collaborative decision-making scenario according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0047] S11, based on the three dimensional indicators of task complexity, knowledge density, and intelligent adaptability of the task to be processed, generate corresponding task features.

[0048] Specifically, task complexity (KL) characterizes the cognitive load, process complexity, and degree of uncertainty required to execute a task. It is calculated based on the number of task steps S, the number of decision points D, and the information dependency I: KL = α1S + β1D + γ1I. Where:

[0049] 1. The number of task steps is S. For example, report preparation requires information collection, data analysis, writing, and review;

[0050] 2. The number of decision points D, if multiple departments need to make collaborative decisions;

[0051] 3. Information dependency (I) is measured from three dimensions: information flow intensity, dependency chain length, and information entropy. Information flow intensity represents the amount of information transferred between statistical tasks (e.g., data volume, call frequency). Dependency chain length indicates that task A depends on the output of task B, which in turn depends on task C, forming a dependency chain. The longer the chain, the higher the I value. Information entropy uses entropy from information theory to measure the uncertainty of dependency links. The higher the entropy, the higher the I value.

[0052] For example, in the report preparation scenario, a simple task (KL=1) refers to the generation of a fixed template report, which only requires filling in data; a complex task (KL=5) refers to the cross-departmental collaborative preparation of an ESG report, which requires integrating multi-source data and meeting regulatory requirements.

[0053] Knowledge density (KD) represents the total amount, expertise, and dispersion of knowledge required to complete a task. It is calculated based on the weighted calculation of three indicators: knowledge base coverage D, expertise depth T, and knowledge relevance C: KD = α2D + β2T + γ2C. Where:

[0054] 1. Knowledge base coverage, which refers to the ratio of the knowledge in the task to be processed to the entire knowledge set in the knowledge base of the corresponding field, such as the ratio of historical cases to be cited in the report;

[0055] 2. Depth of professional knowledge, indicating whether the existing knowledge base alone covers all aspects of knowledge, or whether the experience of industry experts is needed;

[0056] 3. Knowledge relevance indicates the degree of interaction and integration of knowledge between different fields, such as the proportion of knowledge that requires cross-report / cross-system calls. The proportion of cross-field knowledge = the number of knowledge points involving other disciplines / the total number of knowledge points.

[0057] For example, in the report preparation scenario, a low knowledge density task (KD=1) means filling out a daily report in a fixed format; a high knowledge density task (KD=5) means preparing a feasibility report for a new energy project, which requires integrating knowledge from multiple fields such as geology, technology, and finance.

[0058] Intelligent Adaptability (KI) characterizes the ability of intelligent systems (such as RPA and AI models) to support automated tasks. It is calculated based on a weighted combination of three indicators: process automation rate A, algorithm accuracy Z, and human-computer interaction friendliness V: KL = α3A + β3Z + γ3V. Among them:

[0059] 1. Process automation rate is an important indicator for measuring the extent to which technology can replace manual operations in a business process, such as the proportion of steps that can be replaced by RPA.

[0060] 2. Algorithm accuracy is a key indicator for evaluating algorithm performance. It reflects the degree of consistency between the algorithm's prediction results and the actual situation. In the field of natural language processing, this includes the F1 value of the TextCNN error detection model.

[0061] 3. Human-computer interaction friendliness is a comprehensive evaluation of the quality of user experience during the human-computer interaction process, covering multiple dimensions such as interface design, ease of operation, and timely feedback, such as the speed of natural language interaction response.

[0062] For example, in the report preparation scenario, low-fitness tasks (KI=1) require manual verification of data formats one by one; high-fitness tasks (KI=5) use AI to automatically generate a draft report and automatically correct errors through semantic analysis.

[0063] The embodiment of the present invention breaks down each dimension into multiple detailed indicators and performs weighted calculations, thus avoiding the one-sidedness of single indicator evaluation and making task evaluation more in line with actual conditions. Different indicators reflect task characteristics from different dimensions, and through comprehensive consideration in a weighted manner, a more comprehensive and objective evaluation system can be constructed. The weighted calculation method assigns different weights to different indicators, and the weights can be flexibly adjusted according to specific business needs and scenarios. In different industries and at different stages of development, the importance of each indicator to task evaluation may be different. For example, in production-oriented enterprises that emphasize efficiency, the weight of process automation rate in intelligent adaptability evaluation can be appropriately increased; in scientific research institutions that focus on innovation, the weight of professional knowledge depth in knowledge density evaluation should be increased. This flexibility enables the evaluation system to adapt to diverse business needs.

[0064] S12, initialize the initial weight coefficients of the three-dimensional indicators, and use the preset strategy network to dynamically optimize the weight coefficients corresponding to the three-dimensional indicators according to task characteristics and real-time system resource status.

[0065] Specifically, initial weighting factors are determined based on business needs, historical experience, or expert knowledge. For example, in an intelligent customer service system that prioritizes rapid response to customer questions, the initial weighting factor for the "human-computer interaction friendliness" dimension (e.g., natural language interaction response speed) might be set to 0.5. Given the importance of accurately answering questions, the initial weighting factor for the "algorithm accuracy" dimension (e.g., the F1 score of a TextCNN error detection model) might be set to 0.3. Furthermore, the "process automation rate" (e.g., the percentage of manual steps replaced by RPA) plays a significant role in reducing operating costs and is therefore assigned an initial weighting factor of 0.2. These initial values ​​are not fixed but serve as the starting point for subsequent dynamic optimization and will be continuously adjusted based on actual conditions to achieve optimal system performance.

[0066] In one embodiment, based on the trigger conditions for adjusting the preset weight coefficients, the preset strategy network is triggered to dynamically optimize the weight coefficients corresponding to the three dimensional indicators according to the task characteristics and real-time system resource status. The preset trigger conditions are the "switch" for dynamic adjustment of the weight coefficients. Only when these conditions are met will the system start the weight optimization mechanism, avoiding unnecessary frequent adjustments and ensuring the stability and efficiency of the system. In this embodiment of the present invention, the trigger conditions for adjusting the weight coefficients include at least one of the following:

[0067] 1. Time-driven: Weight adjustments are triggered at preset intervals; for example, weights are updated regularly (e.g., quarterly) through the Delphi method. This approach is suitable for business scenarios with cyclical patterns. At each time interval, task execution data, system resource usage, and business indicator completion status are collected. This data is used as the basis for weight coefficient adjustments. Combined with the preset policy network, the weight coefficients are optimized to better match the business rhythm.

[0068] 2. Performance-driven: Weight adjustments are triggered when at least one of the accuracy, efficiency, or cost metrics exceeds a preset threshold. For example, in an intelligent quality inspection system, if the algorithm accuracy falls below 90%, the model is underperforming under the current weight settings, and the weights need to be readjusted to improve detection accuracy. In a production process automation system, if process efficiency drops by 20%, the weights associated with the process automation rate may need to be adjusted to optimize resource allocation and task distribution. These thresholds are typically set based on business objectives, historical data, and industry standards. By monitoring performance metrics in real time and comparing them against thresholds, resource utilization efficiency is improved.

[0069] 3. Task-driven: When the feature distribution of a new task deviates from historical data, a weight adjustment is triggered; for example, the type, complexity, data size, priority and other features of the new task are analyzed and compared with historical task data. When the feature distribution of a new task deviates from historical data to a certain extent, it means that the current weight coefficient may no longer be applicable to the new task, thus triggering an adjustment. For example, in an image recognition system, if the system originally processed images of regular sizes and types, and suddenly received a new task of processing high-definition, complex scene images, the data features will be quite different from the historical data. At this time, the weight needs to be adjusted to optimize the performance of the image recognition algorithm.

[0070] 4. Event-driven: Weight adjustments are triggered when system resource status changes or business priorities shift. For example, changes in system resource status (such as excessive CPU utilization, insufficient memory, or network failures) can impact the system's ability to process tasks, necessitating adjustments to weight coefficients and resource reallocation. For example, when CPU utilization reaches 90%, the weights of metrics requiring high computing resources are reduced to prioritize basic system operations and critical task processing. Business priority changes, such as emergency project launches, alter business priorities and resource requirements, requiring corresponding weight adjustments to ensure the system maintains normal operation in unstable or changing environments and mitigate the negative impact of these events on the business.

[0071] The policy network used in this embodiment is a reinforcement learning model. It learns the optimal policy by continuously interacting with the environment (i.e., selecting actions based on the current state, obtaining rewards, and observing new states). In the weight optimization scenario, "state" is a combination of task characteristics and real-time system resource status, "actions" are adjustments to the weight coefficients of various dimensional indicators, and "rewards" are set based on changes in system performance indicators.

[0072] Before running, the policy network must be trained in a simulated environment or on historical data (such as task completion time, error rate, and labor cost). By learning from a large number of examples, the policy network can understand how to adjust the weight coefficients to achieve optimal system performance under different task characteristics and resource conditions.

[0073] The process of dynamically optimizing the weight coefficients corresponding to the three-dimensional indicators according to the preset strategy network of the embodiment of the present invention based on the task characteristics and real-time system resource status is as follows:

[0074] 1. The task characteristics corresponding to the three-dimensional indicators of the task to be processed and the system resource status are used as the state space. Among them, the task characteristics can be the entropy value Ent(D) corresponding to the three-dimensional indicators, and the system resource status includes the system's business indicators, such as response time and throughput.

[0075] 2. The weight coefficients α, β, and γ corresponding to the three dimensional indicators are used as the action space;

[0076] 3. Use the preset reinforcement learning model to dynamically optimize the weight coefficients corresponding to the three dimensions with the preset reward function. The preset reward function η = accuracy × efficiency - cost, and the weight coefficients α, β, and γ satisfy the constraint α + β + γ = 1.

[0077] This embodiment of the present invention adjusts the weight coefficients (α, β, γ) to reflect the priority of human-machine collaboration in different scenarios:

[0078] - α↑: When the task complexity is high, the weight of human experience increases (such as task decision-making).

[0079] - β↑: When knowledge density is high, the demand for knowledge base and expert intervention increases (such as legal compliance review).

[0080] - γ↑: When the intelligent adaptation degree is high, the automated system takes the lead in execution (such as standardized report generation).

[0081] It should be noted that accuracy in the reward function refers to the probability of correct task completion (for example, when a report is split and assigned to a designated functional department, this is classification accuracy). Efficiency refers to the speed of task execution or resource utilization (for example, the amount of tasks processed per unit time). Cost includes computational overhead, energy consumption, or time delay. The reward function comprehensively considers three important factors: accuracy, efficiency, and cost. Accuracy reflects the quality of task completion, efficiency reflects the speed of task completion, and cost involves resource consumption. This comprehensive evaluation method can comprehensively measure task performance and avoid decision-making bias caused by focusing on a single indicator.

[0082] In an optional embodiment, the reward function incorporates an adjustment factor λ to balance benefits and costs: η = accuracy × efficiency − λ × cost. This adjustment factor can be adjusted based on specific business needs and scenarios. When an enterprise prioritizes benefits, it can be appropriately lowered, shifting the reward function toward improving accuracy and efficiency. When facing significant cost pressures, it can be increased to strengthen cost control. This approach allows for flexible identification of the optimal balance between benefits and costs in different situations, achieving optimal resource allocation.

[0083] S13, weighted normalization is performed on the three-dimensional indicators according to the weight coefficients output by the preset strategy network to generate a comprehensive task score.

[0084] Specifically, the comprehensive score of the task is calculated using the following formula:

[0085]

[0086] The weight coefficient output by the strategy network of the present invention can give different weights to different dimensional indicators according to the characteristics and requirements of the task, thereby highlighting the key factors that have a greater impact on the task. For example, if a task has high requirements for algorithm accuracy, when calculating the comprehensive score, the weight of the dimensional indicator of algorithm accuracy can be increased accordingly, so that the final score can better reflect the performance of the task in key aspects. After weighted normalization of the three dimensional indicators, indicators of different dimensions and different value ranges are converted into a unified scoring scale, and the generated task comprehensive score is comparable and additive. This makes it easy to compare and sort when evaluating multiple tasks, which helps decision makers quickly screen out tasks with higher priority or better performance based on the comprehensive score, and provides strong support for decisions such as resource allocation and task scheduling.

[0087] S14, dividing the task priorities according to the comprehensive task scores, and dynamically allocating the tasks to manual and / or intelligent systems based on the priorities to perform the tasks.

[0088] In an optional embodiment, the comprehensive score of the task is divided into three threshold intervals according to the business needs or historical data of the task to be processed. Different threshold intervals correspond to different priorities. According to the mapping relationship between the preset priority and the task allocation and the threshold interval in which the comprehensive score of the task to be processed is located, it is confirmed that the task is executed by any of the manual, intelligent system, and human-computer collaboration methods.

[0089] Specifically, based on business needs or historical data analysis, three threshold intervals are set (the thresholds can be dynamically adjusted through cluster analysis or ROC curves). Different threshold intervals belong to the three priority categories of High / Medium / Low, for example:

[0090] High: P ≥ 0.7;

[0091] Medium: 0.4≤P<0.7;

[0092] Low: P < 0.4

[0093] The mapping relationship between preset priorities and task allocation, for example, High priority is assigned to manual processing, Medium priority is assigned to human-machine collaborative processing, and Low priority is assigned to intelligent systems. This is for example only and is not limited to this. By obtaining the comprehensive score of the pending tasks, the priority level is determined and the corresponding resource allocation method is further matched.

[0094] In a practical application, when the backlog of High tasks exceeds the preset time (for example, 3 hours), it will trigger the expansion of human resources, such as the temporary deployment of other personnel; by confirming the priority of task complexity, the optimization method of deploying resources such as people, data, and computing power is used to maximize resource utilization efficiency (people are put to the best use, data is accurately matched, and computing power is allocated on demand); achieve optimization of productivity structure (improvement of human-machine collaboration efficiency, and balance of standardization and agility); achieve dual control of costs and risks (cost optimization, risk prevention and control); and achieve organizational effectiveness upgrades (agile response capabilities, capability accumulation and innovation. For example, automation experience of low-complexity tasks can be accumulated as standard tools, such as the RPA robot library, to reduce duplication of construction).

[0095] In another alternative implementation, when there are multiple pending tasks, their comprehensive scores are compared and the task allocation method is determined based on the comparison results. Each task is assigned to either manual execution, an intelligent system, or human-machine collaboration. This resource allocation method aligns with the "efficiency first" principle of resource management, avoiding resource waste or imbalanced allocation. It further achieves the rational scheduling of resources and the optimal allocation of productivity factors, focusing on factors such as people, data, and computing power.

[0096] For example, in a report preparation task assignment scenario:

[0097] Input data:

[0098] Task A (prepare monthly operating report): KL=3, KD=4, KI=2;

[0099] Task B (preparing an annual ESG report): KL=5, KD=5, KI=4;

[0100] Weight setting (focusing on knowledge density): α=0.2, β=0.5, γ=0.3;

[0101] Calculation results:

[0102] f(TC)_A = 0.2×3 + 0.5×4 + 0.3×2 = 3.2;

[0103] f(TC)_B = 0.2×5 + 0.5×5 + 0.3×4 = 4.7;

[0104] decision making:

[0105] Task B has a higher TC and prioritizes knowledge density, and is assigned to an expert team to lead, with AI assisting in data verification.

[0106] For example, in the intelligent customer service system optimization scenario:

[0107] Input data:

[0108] User consultation type C (product troubleshooting): KL=4, KD=3, KI=5;

[0109] User inquiry type D (package price inquiry): KL=2, KD=1, KI=4;

[0110] Weight setting (focusing on intelligent adaptation): α=0.1, β=0.2, γ=0.7;

[0111] Calculation results:

[0112] f(TC)_C = 0.1×4 + 0.2×3 + 0.7×5 = 4.1;

[0113] f(TC)_D = 0.1×2 + 0.2×1 + 0.7×4 = 3.2;

[0114] Decision: Type C has a higher TC and high intelligent adaptability, so AI self-service is enabled first; Type D is transferred to manual customer service.

[0115] The embodiments of the present invention are also applicable to different industries, such as energy and medical care. For example, in the energy industry, in the power equipment maintenance task, the knowledge density KD calculation needs to additionally introduce the "equipment model compatibility" parameter, KD=0.5D+0.2T+0.1C+0.1M, where M is the equipment model matching degree, which is calculated through the total compatibility labels of historical maintenance records.

[0116] This embodiment of the present invention prioritizes tasks scientifically, ensuring that important and urgent tasks are handled first, avoiding critical business delays caused by improper task processing order. It also rationally allocates tasks to both manual and intelligent systems, leveraging the strengths of both, reducing task waiting times and processing times, and improving the overall speed of business processes.

[0117] By dynamically allocating tasks to manual and intelligent systems, the present invention allows for flexible adjustments based on resource load, avoiding idle or overused resources. This rationally arranges manual work to avoid wasted manpower and overwork. By fully leveraging the automated processing capabilities of intelligent systems, the system improves versatility across multiple task scenarios, significantly enhancing the dynamic adaptability, decision-making reliability, and task execution efficiency of human-machine collaborative systems, improving computing resource utilization, and reducing enterprise operating costs.

[0118] To verify the effectiveness of the resource allocation method provided by the embodiment of the present invention in the human-machine collaborative decision-making scenario, an energy company used a model to optimize the report preparation process. The results of the traditional model and the human-machine collaborative model are compared as shown in Table 1 below:

[0119] Table 1

[0120]

[0121] Through the above comparison, it is found that the human-machine collaborative mode significantly improves efficiency, verifying the effectiveness of the method provided by the embodiment of the present invention.

[0122] The embodiment of the present invention also provides a resource allocation method in a human-machine collaborative decision-making scenario, which introduces the three-dimensional indicators of task complexity, knowledge density, and intelligent adaptability based on the task to be processed.

[0123] Time sensitivity (TS) and resource load (RL) serve as moderation factors.

[0124] Specifically, such as Figure 2 As shown, the following steps are included:

[0125] S21 generates task features based on the five dimensional indicators of task complexity, knowledge density, intelligent adaptability, time sensitivity and resource load of the task to be processed.

[0126] Specifically, time sensitivity (TS) measures task sensitivity based on its urgency; resource load (RL) is the resource load of the node to which the task is currently assigned. If the resource load has a nonlinear relationship with the original state variable (for example, the resource load is a dynamically changing bottleneck threshold), the introduction of these two adjustment factors allows the task allocation model to comprehensively consider more factors, no longer limited to the characteristics of the task itself, but also including the task's time requirements and the system's resource status. This makes task allocation more realistic and can more accurately assign tasks to the most suitable artificial or intelligent systems for processing, improving the success rate and quality of task processing. For example, in software development projects, tasks with high time sensitivity and knowledge density will be assigned to experienced developers, taking into account their current resource load to ensure that they have sufficient energy and resources to complete the task with high quality.

[0127] S22: Use the analytic hierarchy process or entropy weight method to initialize the initial weight coefficients of the five dimensions of task complexity, knowledge density, intelligent adaptability, time sensitivity, and resource load of the task to be processed, and build a real-time feedback loop to optimize the weight distribution through the preset strategy network;

[0128] Specifically, the hierarchical analysis method is suitable for scenarios that require expert experience. It can weight all factors (KL, KD, KI, TS, RL), establish a hierarchical structure based on expert experience and task importance, and provide explainable business logic. The entropy weight method is suitable for purely data-driven scenarios. Based on the entropy of data information, it only weights factors with historical data and objectively reflects the variability of indicators to determine weights. Both methods can be flexibly selected according to actual needs to ensure that the initial weights conform to business logic and data characteristics.

[0129] By building a real-time feedback loop, reinforcement learning models can dynamically adjust weights based on task execution results to achieve adaptive optimization. Faced with complex and uncertain task scenarios, using reinforcement learning models through real-time feedback loops to adjust weights based on task execution results can better cope with various emergencies and uncertainties. For example, in network security defense, the model dynamically adjusts defense strategy weights based on the outcome of attack incidents, promptly strengthening protection of weak links, reducing the risk of system attacks, and improving the system's overall anti-interference capabilities and stability.

[0130] S23, weighted normalization of the five dimension indicators is performed according to the weight coefficients output by the preset strategy network to generate a comprehensive task score;

[0131] S24, dividing the task priorities according to the comprehensive task scores, and dynamically allocating the tasks to manual and / or intelligent systems based on the priorities to perform the tasks.

[0132] The processes of steps S23 and S24 are similar to those of steps S13 and S14 respectively, and are not described again here.

[0133] The embodiment of the present invention combines time sensitivity and resource load with task complexity, knowledge density, and intelligent adaptability, comprehensively covering task attributes and execution environment factors, helping the system to more accurately grasp the overall picture of the task, provide a more reliable basis for task allocation, and ensure that the task allocation strategy can simultaneously meet task characteristics and system resource conditions. The real-time feedback loop and dynamic weight adjustment mechanism enable the system to quickly adapt to task requirements and environmental changes. Whether it is a change in task priority, resource fluctuations, or the addition of new tasks, the system can quickly adjust the task allocation strategy to maintain efficient operation and enhance the system's adaptability and anti-interference capabilities in complex environments.

[0134] In this embodiment, a resource allocation system for a human-machine collaborative decision-making scenario is also provided. The system is used to implement the above-mentioned embodiments and preferred implementation modes, and the details that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the system described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0135] This embodiment provides a resource allocation system in a human-machine collaborative decision-making scenario, such as Figure 3 Shown, including:

[0136] A three-dimensional evaluation model building module 31 is used to generate corresponding task features based on the three dimensional indicators of task complexity, knowledge density, and intelligent adaptability of the task to be processed;

[0137] The dynamic weight adjustment module 32 is used to initialize the initial weight coefficients of the three-dimensional indicators and dynamically optimize the weight coefficients corresponding to the three-dimensional indicators based on the task characteristics and real-time system resource status using the preset strategy network, including:

[0138] The task characteristics and system resource status corresponding to the three dimensional indicators of the task to be processed are used as the state space;

[0139] The weight coefficients α, β, and γ corresponding to the three dimensional indicators are used as the action space;

[0140] Using a preset reinforcement learning model, the weight coefficients corresponding to the three dimensions are dynamically optimized with a preset reward function, where the preset reward function η = accuracy × efficiency - cost, and the weight coefficients α, β, and γ satisfy the constraint α + β + γ = 1;

[0141] The task comprehensive score calculation module 33 is used to perform weighted normalization on the three dimensional indicators according to the weight coefficients output by the preset strategy network to generate a task comprehensive score;

[0142] The task allocation module 34 is used to prioritize tasks according to their comprehensive scores, and dynamically allocate tasks to manual and / or intelligent systems based on their priorities to perform the tasks.

[0143] In some optional embodiments, the task complexity in the three-dimensional evaluation model construction module 31 is obtained by weighted calculation based on the three indicators of the number of task steps, decision point density, and information dependency; the knowledge density is obtained by weighted calculation based on the three indicators of knowledge base coverage, professional knowledge depth, and knowledge relevance; and the intelligent adaptability is obtained by weighted calculation based on the three indicators of process automation rate, algorithm accuracy, and human-computer interaction friendliness.

[0144] In some optional embodiments, the dynamic weight adjustment module 32 triggers the preset policy network based on the trigger conditions of the preset weight coefficient adjustment according to the task characteristics and the real-time system resource status, and the trigger conditions of the weight coefficient adjustment include at least one of the following time-driven, triggering the weight adjustment at a preset time interval;

[0145] Performance-driven, weight adjustment is triggered when at least one of the accuracy, efficiency or cost indicators exceeds a preset threshold;

[0146] Task-driven, when the feature distribution of a new task deviates from historical data, it triggers weight adjustment;

[0147] Event-driven, weight adjustment is triggered when the system resource status changes or business priorities change.

[0148] In some optional implementations, the task assignment module 34 includes:

[0149] The first allocation unit divides the comprehensive score of the task into three threshold intervals based on the business needs or historical data of the task to be processed. Different threshold intervals correspond to different priorities. Based on the mapping relationship between the preset priority and the task allocation and the threshold interval in which the comprehensive score of the task to be processed falls, the first allocation unit determines whether to execute the task manually, by an intelligent system, or by human-machine collaboration.

[0150] The second allocation unit is used to compare the comprehensive scores of the tasks to be processed when there are multiple tasks to be processed, and determine the task allocation method according to the comparison result.

[0151] In some optional embodiments, the method further includes:

[0152] The adjustment factor introduction module is used to generate task features based on the time sensitivity and resource load indicators of the tasks to be processed;

[0153] The real-time weight optimization module is used to initialize the initial weight coefficients of the five dimensions of task complexity, knowledge density, intelligent adaptability, time sensitivity and resource load of the task to be processed using the hierarchical analysis method or entropy weight method, and build a real-time feedback loop to optimize the weight distribution through a preset strategy network.

[0154] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0155] The resource allocation system in the human-machine collaborative decision-making scenario in this embodiment is presented in the form of functional units, where the units refer to ASIC (Application Specific Integrated Circuit) circuits, processors and memories that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0156] The embodiment of the present invention also provides a computer device having the above Figure 3 The resource allocation system in the human-machine collaborative decision-making scenario shown.

[0157] See also Figure 4 , Figure 4 Schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 4 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 A processor 10 is taken as an example.

[0158] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0159] The memory 20 stores instructions that can be executed by at least one processor 10, so as to enable at least one processor 10 to execute the method shown in the above embodiment.

[0160] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0161] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0162] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0163] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor central control system or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0164] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0165] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A resource allocation method in a human-machine collaborative decision-making scenario, characterized in that: include: Based on the three dimensional indicators of task complexity, knowledge density, and intelligent adaptability of the task to be processed, the corresponding task features are generated. The task complexity is calculated by weighting the three indicators of task steps, decision point density, and information dependency; the knowledge density is calculated by weighting the three indicators of knowledge base coverage, professional knowledge depth, and knowledge relevance; and the intelligent adaptability is calculated by weighting the three indicators of process automation rate, algorithm accuracy, and human-computer interaction friendliness. Initialize the initial weight coefficients of the three-dimensional indicators, and use the preset strategy network to dynamically optimize the weight coefficients corresponding to the three-dimensional indicators according to task characteristics and real-time system resource status, including: The task characteristics and system resource status corresponding to the three dimensional indicators of the task to be processed are used as the state space; The weight coefficients α, β, and γ corresponding to the three dimensional indicators are used as the action space; Using a preset reinforcement learning model, the weight coefficients corresponding to the three dimensions are dynamically optimized with a preset reward function, where the preset reward function η = accuracy × efficiency - cost, and the weight coefficients α, β, and γ satisfy the constraint α + β + γ = 1; The three-dimensional indicators are weighted and normalized according to the weight coefficients output by the preset strategy network to generate a comprehensive task score; Tasks are prioritized based on their comprehensive scores and dynamically assigned to manual and / or intelligent systems for execution based on priority.

2. The method according to claim 1, characterized in that Based on the triggering conditions for adjusting the preset weight coefficients, the preset policy network is triggered according to the task characteristics and the real-time system resource status, and the triggering conditions for adjusting the weight coefficients include at least one of the following: Time-driven, triggering weight adjustments at preset time intervals; Performance-driven, weight adjustment is triggered when at least one of the accuracy, efficiency or cost indicators exceeds a preset threshold; Task-driven, when the feature distribution of a new task deviates from historical data, it triggers weight adjustment; Event-driven, weight adjustment is triggered when the system resource status changes or business priorities change.

3. The method according to claim 1, characterized in that The method of prioritizing tasks according to their comprehensive scores and dynamically allocating tasks to manual and / or intelligent systems based on their priorities includes: Divide the task's comprehensive score into three threshold intervals based on the business needs or historical data of the task to be processed. Different threshold intervals correspond to different priorities. Based on the mapping relationship between the preset priority and task allocation and the threshold interval in which the task's comprehensive score falls, determine whether to execute the task manually, using an intelligent system, or using human-machine collaboration; or When there are multiple tasks to be processed, the comprehensive scores of the tasks to be processed are compared, and the task allocation method is determined based on the comparison results.

4. The method according to claim 3, characterized in that Also includes: Generate task features based on the time sensitivity and resource load indicators of the tasks to be processed; The initial weight coefficients of the five dimensions of task complexity, knowledge density, intelligent adaptability, time sensitivity and resource load of the task to be processed are initialized using the hierarchical analysis method or entropy weight method, and a real-time feedback loop is constructed to dynamically optimize the weight coefficients through a preset strategy network.

5. A resource allocation system in a human-machine collaborative decision-making scenario, characterized in that: include: A three-dimensional evaluation model construction module is used to generate corresponding task features based on the three dimensional indicators of task complexity, knowledge density, and intelligent adaptability of the task to be processed. The task complexity is calculated by weighting the three indicators of the number of task steps, decision point density, and information dependency; the knowledge density is calculated by weighting the three indicators of knowledge base coverage, professional knowledge depth, and knowledge relevance; and the intelligent adaptability is calculated by weighting the three indicators of process automation rate, algorithm accuracy, and human-computer interaction friendliness. The dynamic weight adjustment module is used to initialize the initial weight coefficients of the three-dimensional indicators and dynamically optimize the weight coefficients corresponding to the three-dimensional indicators based on the preset strategy network according to the task characteristics and real-time system resource status; The task comprehensive score calculation module is used to weight and normalize the three-dimensional indicators according to the weight coefficients output by the preset strategy network to generate a comprehensive task score, including: The task characteristics and system resource status corresponding to the three dimensional indicators of the task to be processed are used as the state space; The weight coefficients α, β, and γ corresponding to the three dimensional indicators are used as the action space; Using a preset reinforcement learning model, the weight coefficients corresponding to the three dimensions are dynamically optimized with a preset reward function, where the preset reward function η = accuracy × efficiency - cost, and the weight coefficients α, β, and γ satisfy the constraint α + β + γ = 1; The task allocation module is used to prioritize tasks based on their comprehensive scores and dynamically allocate tasks to manual and / or intelligent systems based on their priorities.

6. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the resource allocation method in the human-computer collaborative decision-making scenario according to any one of claims 1 to 4 by executing the computer instructions.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the resource allocation method in the human-computer collaborative decision-making scenario according to any one of claims 1 to 4.

8. A computer program product, characterized in that It includes computer instructions, which are used to enable a computer to execute the resource allocation method in the human-computer collaborative decision-making scenario described in any one of claims 1-4.

Citation Information

Patent Citations

  • Cooperative interaction intelligent access method of all-in-one computer

    CN120011089A