Digital labor system closed-loop regulation and control method based on multi-agent reinforcement learning

By employing a multi-agent reinforcement learning approach, we can perceive and regulate multi-source data in a digital workforce system in real time, construct a multi-party Markov game model, and generate stability-oriented regulation strategies. This approach solves the problems of disconnect between perception and decision-making and the evolution of group behavior, thereby improving the stability and robustness of the system.

CN121961139APending Publication Date: 2026-05-01QINGDAO UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO UNIV OF TECH
Filing Date
2026-01-23
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing multi-agent reinforcement learning frameworks in digital workforce systems suffer from problems such as a disconnect between perception and decision-making, insufficient modeling of group game theory, and a lack of stability objectives. They are unable to effectively cope with changes in individual worker behavior and the evolution of group behavior, thus affecting system stability and robustness.

Method used

By employing a multi-agent reinforcement learning approach, workers, platform providers, and regulatory stakeholders are modeled as multiple interacting agents. Through multi-source data perception and state construction, multi-agent game modeling, reinforcement learning decision generation, and stability assessment, a closed-loop control mechanism is formed to achieve real-time perception, prediction, and control of multi-party game behavior.

Benefits of technology

It enhances the overall stability and robustness of the digital workforce system, and through real-time data perception and strategy adjustment, it strengthens the ability to model and predict group collaborative or adversarial behaviors, ensuring the stable operation of the system in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961139A_ABST
    Figure CN121961139A_ABST
Patent Text Reader

Abstract

The invention relates to the cross technical field of artificial intelligence and digital platform governance, and discloses a digital labor system closed-loop regulation and control method based on multi-agent reinforcement learning, which comprises the following steps: S1, collecting multi-source data in real time and fusing, and constructing a system state vector; s2, modeling key participants of the system into four types of intelligent agents including a platform, a worker, a demand side and a constraint, and constructing a multi-party Markov game model; s3, a centralized reviewer and distributed actuator architecture is adopted, and a combined regulation and control strategy with stability as a target is generated; s4, calculating three types of stability indexes of the laborer, the demand side and the whole system, and mapping the three types of stability indexes into reinforcement learning feedback signals; and S5, mapping the strategy into a platform executable parameter, performing execution after security verification, and feeding back an execution result to the sensing layer to form a closed-loop regulation and control mechanism. According to the method, multi-agent reinforcement learning is adopted to model a multi-party game, real-time sensing, prediction and regulation are realized, and the stability and robustness of a labor system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

A Closed-Loop Regulation Method for Digital Labor Force Systems Based on Multi-Agent Reinforcement Learning Technical Field

[0001] This application belongs to the interdisciplinary field of artificial intelligence and digital platform governance, and specifically relates to a closed-loop control method for digital workforce systems based on multi-agent reinforcement learning. Background Technology

[0002] A digital labor system, also known as a gig economy platform or on-demand service platform, is a new type of labor organization and resource allocation system built on the internet and digital technologies. Its key characteristic is the dynamic connection of a massive number of dispersed service providers (such as ride-hailing drivers, food delivery riders, and freelancers) with service demanders (such as passengers, consumers, and businesses) through a centralized algorithm platform, enabling fully online and automated management of the entire process of task requests, labor scheduling, payroll settlement, and service evaluation.

[0003] Against the backdrop of the rapid development of digital platforms and the gig economy, new employment forms such as ride-hailing, food delivery, and digital content creation are highly dependent on algorithm-driven resource allocation and incentive mechanisms. The stability of the labor force system directly affects platform operational efficiency, worker rights protection, and social governance costs, and has become a key technical issue in digital platform governance. While existing platform control technologies have made some progress in addressing changes in individual worker behavior, they still have significant shortcomings in multi-agent interaction, group behavior evolution, and the long-term maintenance of system stability.

[0004] Existing methods include behavioral analysis methods based on evolutionary game theory, dynamic incentive and scheduling methods based on reinforcement learning, and general game frameworks based on multi-agent reinforcement learning. However, existing multi-agent reinforcement learning frameworks mainly focus on optimizing competitive or collaborative efficiency and have the following shortcomings: Disconnect between perception and decision-making: Workers' psychological perceptions and behavioral intentions are difficult to map into executable algorithm control parameters in real time, lacking a closed-loop mechanism between perception and policy; Insufficient group game modeling: Failing to effectively address adversarial or collaborative group behaviors among workers, resulting in a lack of system stability resilience; Missing stability objectives: Existing multi-agent methods emphasize efficiency or profit maximization, lacking clear modeling and optimization objectives for labor system stability indicators.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] To address or at least alleviate one or more of the above problems, a closed-loop control method for digital labor force systems based on multi-agent reinforcement learning is provided. By modeling workers, platform providers, and regulatory stakeholders as multiple interacting agents and introducing a multi-agent reinforcement learning mechanism, the method enables real-time perception, prediction, and control of multi-party game behavior in a dynamic environment, thereby improving the overall stability and robustness of the labor force system.

[0007] To achieve the above objectives, according to the first aspect of this application, a closed-loop control method for a digital workforce system based on multi-agent reinforcement learning is provided, comprising: S1, multi-source data perception and state construction: real-time acquisition of multi-source data from the digital workforce system, preprocessing and fusion, and construction of a unified system state vector. The multi-source data includes at least labor supply-side data, demand-side data, platform regulation data, and system constraint and historical feedback data; S2, Multi-agent game modeling: The key participants in the digital labor system are modeled as four group-level agents, namely, platform agents. Labor group intelligence Demand-side collective intelligence and system-constrained intelligent agents Furthermore, a multi-party Markov game model is constructed to describe the dynamic transition of system states and the interaction of multi-party strategies; S3, Reinforcement Learning Decision Generation: A multi-agent reinforcement learning architecture combining a centralized commentator and a distributed executor is adopted, based on the system state vector. A joint control strategy is generated, with the system stability assessment result as the optimization objective; S4, Stability assessment and feedback signal generation: Based on the stability index of the worker group, the demand-side operation stability index, and the overall system coordination stability index, the system stability level is calculated and mapped to reinforcement learning real-time feedback signals. S5, Closed-loop control execution and feedback: The joint control strategy output by reinforcement learning is mapped to platform executable parameters. Security verification is performed before execution, and the system operation results after control are fed back to multi-source data perception in real time, forming a closed-loop control mechanism from S1 to S5.

[0008] To achieve the above objectives, according to the second aspect of this application, a closed-loop control system for a digital workforce system based on multi-agent reinforcement learning is provided. The closed-loop control system includes: a multi-source data perception layer, used to collect multi-source data from the digital workforce system in real time, perform preprocessing and fusion, and construct a unified system state vector. The multi-source data includes at least labor supply-side data, demand-side data, platform regulation data, and system constraint and historical feedback data. The multi-agent modeling layer is used to model the key participants in the digital workforce system as four group-level agents, namely, platform agents. Labor group intelligence Demand-side collective intelligence and system-constrained intelligent agents A multi-party Markov game model is constructed to describe the dynamic transition of system states and the interaction of multi-party strategies; a reinforcement learning decision layer is used to employ a multi-agent reinforcement learning architecture that combines a centralized commentator and distributed executors, based on the system state vector. A joint control strategy is generated, with the system stability assessment result as the optimization objective. A stability assessment layer is used to calculate the system stability level based on worker group stability indicators, demand-side operational stability indicators, and overall system coordination stability indicators, and maps this level to reinforcement learning real-time feedback signals. The closed-loop control feedback layer is used to map the joint control strategy output by reinforcement learning into platform executable parameters, perform security verification before execution, and feed back the system operation results after control to the multi-source data perception in real time, forming a closed-loop control mechanism from the multi-source data perception layer to the closed-loop control execution and feedback layer.

[0009] To achieve the above objectives, in accordance with a third aspect of this application, a computer-readable storage medium is provided, storing a computer program that, when executed by a processor, is used to implement the closed-loop control method for a digital workforce system based on multi-agent reinforcement learning as described above.

[0010] After adopting the above technical solution, this application has the following beneficial effects compared with the prior art: Through multi-source data perception and state construction, this application enables the system to collect and fuse multi-dimensional data (supply side, demand side, platform regulation, system constraints) in real time, and construct a unified system state vector. It achieves real-time, structured perception of the system's operational status; combined with S3 reinforcement learning decision generation and S5 closed-loop feedback, it forms an end-to-end closed loop from perception to decision to execution to feedback, enabling perception data to directly drive policy adjustments and solving the problem of disconnect between perception and decision.

[0011] This application models the platform, workers, demand side, and system constraints as a group-level intelligent agent through multi-agent game modeling, and constructs a multi-party Markov game model to explicitly characterize the multi-party strategy interaction and state transition. Combined with the centralized commentator and distributed executor architecture in S3, the system can collaboratively learn multi-party strategies, enhance the modeling and prediction capabilities of group collaborative or adversarial behavior, and improve the system's game resilience and adaptability.

[0012] This application constructs a multi-dimensional stability index system (labor stability, demand-side stability, and system coordination stability) through stability assessment and feedback signal generation, and maps the stability assessment results into reinforcement learning reward signals. The system is based on To optimize the objective, a strategy search is conducted, ensuring that the decision-making process always revolves around system stability. This achieves stability-oriented adaptive control and avoids the risk of system instability caused by optimizing a single objective.

[0013] This application employs a feedback mechanism to create a dual-loop structure (operational loop and strategy learning loop), supporting continuous online learning and strategy correction to adapt to environmental non-stationarity and behavioral evolution. It adopts a modular, layered architecture, with each layer connected through standardized interfaces, facilitating deployment and expansion across different digital workforce platforms. Decision-making latency is controllable, supporting rapid response in dynamic markets, and security checks and embedded constraints ensure a safe, compliant, and stable regulatory process.

[0014] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. Attached Figure Description

[0015] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application. The illustrative embodiments and descriptions of the application are used to explain the application, but do not constitute an undue limitation of the application. Obviously, the drawings described below are merely some embodiments, and those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0016] In the accompanying drawings: Figure 1 is a schematic diagram of the architecture of the closed-loop control system of the digital labor force system based on multi-agent reinforcement learning in this specific embodiment; Figure 2 is a schematic diagram of the constraint effect of the system constraint agent on the joint policy space in this specific embodiment; Figure 3 is a schematic diagram of the multi-agent joint decision-making structure under stability evaluation constraints in this specific embodiment; Figure 4 is a flowchart of the stability evaluation layer structure and feedback signal generation in this specific embodiment; Figure 5 is a flowchart of the dual closed-loop control mechanism in this specific embodiment. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.

[0018] Example 1: Please refer to Figure 1. This application provides a closed-loop control system for a digital labor force system based on multi-agent reinforcement learning. The closed-loop control system for the digital labor force system includes: a multi-source data perception layer, used to collect multi-source data from the digital labor force system in real time. The multi-source data includes at least labor supply-side data, demand-side data, platform control data, and system constraint and historical feedback data. The data is preprocessed and fused to construct a unified system state vector. The multi-agent modeling layer is used to model the key stakeholders in the system as four group-level agents, including the platform agent. Labor group intelligence Demand-side collective intelligence and system-constrained intelligent agents A multi-party Markov game model is constructed to describe the dynamic transition of system states and the interaction of multi-party strategies; a reinforcement learning decision layer is used to employ a multi-agent reinforcement learning architecture that combines a centralized commentator and distributed executors, based on the system state. A joint control strategy is generated, with the system stability assessment result as the optimization objective. A stability assessment layer is used to calculate the system stability level based on worker group stability indicators, demand-side operational stability indicators, and overall system coordination stability indicators, and maps this level to reinforcement learning real-time feedback signals. The closed-loop control feedback layer is used to map the control strategy output by reinforcement learning into platform executable parameters, perform security verification before execution, and feed back the system operation results after control to the multi-source data perception layer in real time, forming a closed-loop control mechanism.

[0019] Multi-source data perception layer: This layer is used for comprehensive and continuous perception and collection of the operational status of the digital workforce system. It receives data from multi-dimensional information sources during platform operation, reflecting changes in labor supply and demand, platform control behaviors, and historical system feedback, providing fundamental state input for subsequent agent modeling and decision-making. The perception layer outputs system state descriptions to the upper layers using standardized data interfaces, ensuring the real-time nature, completeness, and scalability of the data.

[0020] Multi-agent Modeling Layer: This layer is used to abstractly model key stakeholders in the digital workforce system. Based on stakeholder game theory, it models the platform, the workforce, the demand-side environment, and system constraints at the role level, clarifying the state characteristics, optional behaviors, and mutual influence relationships of each participant, thus forming a multi-party interactive system model. Through this layer, the system can depict the strategic game relationships between different stakeholders in a dynamic environment, providing a structured foundation for decision optimization.

[0021] Reinforcement learning decision layer: Used to generate dynamic control strategies in multi-agent interactive environments. This layer, within a unified decision framework, jointly optimizes the behavior of multiple agents and outputs corresponding control decisions based on system state changes, guiding the dynamic adjustment of platform operating parameters and strategy configurations. The decision layer supports online learning and strategy updates to adapt to the non-stationarity and uncertainty in the operation of the digital workforce system.

[0022] Stability Assessment Layer: This layer comprehensively evaluates the operational status of the digital workforce system. Based on system operation results and feedback information, it quantitatively assesses the system's stability, coordination, and risk level, and transforms the assessment results into decision feedback signals to guide strategy optimization. The stability assessment results serve as a crucial basis for closed-loop system control, ensuring that control objectives always revolve around the overall stable operation of the system.

[0023] Closed-loop control feedback layer: This layer applies the control strategies output by the decision-making layer to the actual platform operating system. It dynamically adjusts platform control parameters, resource allocation methods, or operating rules to intervene in and guide the system's operating state. Simultaneously, the control execution results are fed back to the data perception layer in real time, forming a complete closed-loop control process, thereby supporting continuous adaptive optimization of the system.

[0024] The closed-loop control method for a digital workforce system based on multi-agent reinforcement learning provided in this application includes: S1, multi-source data perception and state construction: real-time acquisition of multi-source data from the digital workforce system, preprocessing and fusion, and construction of a unified system state vector. The multi-source data includes at least labor supply-side data, demand-side data, platform regulation data, and system constraint and historical feedback data; S2, Multi-agent game modeling: The key participants in the digital labor system are modeled as four group-level agents, namely, platform agents. Labor group intelligence Demand-side collective intelligence and system-constrained intelligent agents Furthermore, a multi-party Markov game model is constructed to describe the dynamic transition of system states and the interaction of multi-party strategies; S3, Reinforcement Learning Decision Generation: A multi-agent reinforcement learning architecture combining a centralized commentator and a distributed executor is adopted, based on the system state vector. A joint control strategy is generated, with the system stability assessment result as the optimization objective; S4, Stability assessment and feedback signal generation: Based on the stability index of the worker group, the demand-side operation stability index, and the overall system coordination stability index, the system stability level is calculated and mapped to reinforcement learning real-time feedback signals. S5, Closed-loop control execution and feedback: The joint control strategy output by reinforcement learning is mapped to platform executable parameters. Security verification is performed before execution, and the system operation results after control are fed back to multi-source data perception in real time, forming a closed-loop control mechanism from S1 to S5.

[0025] It should be noted that the closed-loop control method for digital labor force system based on multi-agent reinforcement learning provided in this application is implemented through the above layers, and the specific method steps are explained in conjunction with each layer.

[0026] In some embodiments, the multi-source data perception layer is used to uniformly and continuously perceive and aggregate the supply side, demand side, platform side, and system constraint state of the digital workforce system, providing a state basis for multi-agent game modeling and reinforcement learning decision-making.

[0027] Sensing objects and data sources: The multi-source data sensing layer includes at least the following four types of data collection modules: The labor supply side data collection module is used to collect data reflecting the operational status of the labor group, including but not limited to: group order acceptance rate, online rate and task completion rate; average income and fluctuation range; order rejection, offline or abnormal behavior ratio; historical retention and turnover trend characteristics.

[0028] The demand-side data collection module is used to collect data reflecting the demand behavior of users or employers, including but not limited to: order generation intensity and time distribution; sensitivity to price, waiting time and service quality; demand elasticity and volatility characteristics; historical demand stability and records of sudden changes.

[0029] The platform control data acquisition module is used to collect control and operation parameters on the platform side, including but not limited to: order dispatch rules and priority parameters; subsidy or incentive intensity range; rule adjustment frequency and magnitude.

[0030] The system constraint and historical feedback data module is used to collect historical assessment results of system stability, compliance constraint information, and risk event records.

[0031] Unified State Representation: After denoising, normalizing, and time-aligning the multi-source data, it is mapped to a unified system state vector. , is represented as: ;in, Indicates the characteristics of the working group's status. This indicates the characteristics of the demand-side group's state; Indicates the characteristics of the platform's control status; This represents the system constraints and historical feedback characteristics.

[0032] This state vector serves as the input for subsequent multi-agent game modeling and reinforcement learning decision-making, enabling a holistic understanding of the system's operational status.

[0033] In some embodiments, the multi-agent modeling layer employs a stakeholder-based group-level modeling approach to abstract key stakeholders with strategic influence in the digital workforce system into four types of strategic agents.

[0034] Definition of an intelligent agent set: The key participants in a digital workforce system are abstracted into an intelligent agent set consisting of the following four group-level intelligent agents, represented as: ;in, It serves as the platform's intelligent agent, used to output platform-side control actions, including dynamic subsidy coefficients, order assignment weights, or rule tightness parameters. It is a collective intelligence agent for workers, used to characterize the behavioral response tendencies of workers as a whole or in subgroups, including the probability of order acceptance or the trend of activity changes. It is a demand-side collective intelligence agent used to describe the demand elasticity and behavioral feedback of users or employers under changes in price or service quality. For system-constrained agents, it is used to introduce stability, security or compliance constraints that must be met during system regulation. By influencing the system state transition and the joint strategy feasible domain, the constraints are embedded in the multi-agent game process.

[0035] Please refer to Figure 2. Figure 2 illustrates how the system-constrained agent operates on the joint policy space. By dynamically constraining the feasible domain of the joint policy, the system-constrained agent embeds compliance and security requirements into the multi-agent game and policy optimization process. Instead of directly participating in the payoff game, it acts as a "rule guardian," dynamically defining the feasible boundaries of joint strategies and identifying high-risk areas based on real-time assessments of system stability, compliance requirements, and historical risks. This process seamlessly embeds hard and soft constraints into the multi-agent interaction model, ensuring that all strategy searches and optimizations in the reinforcement learning decision layer are automatically confined to safe and compliant feasible domains, thereby fundamentally guaranteeing the long-term stability of the system.

[0036] Each intelligent agent is a strategy abstraction of the group behavior with similar decision-making characteristics.

[0037] Multi-Agent Markov Game Modeling: The system is modeled as a multi-agent Markov game, and its state transition function is defined as: Where t represents the discrete decision time step, This represents the global state of the system at time step t, output by the multi-source data sensing layer. This indicates the system state at the next moment after the control action is performed. Indicates platform intelligent agent The regulatory actions taken at time step t Represents the collective intelligence of the working class The action taken at time step t. Represents the demand-side collective intelligence agent The demand response actions taken at time step t Represents the system's constrained intelligent agent The constraint action or constraint condition applied at time step t This represents the system state transition probability function, used to describe the evolution of the system state under the joint actions of multiple agents.

[0038] This modeling approach enables the depiction of the dynamic game-theoretic relationship between platform, supply, demand, and constraints.

[0039] In some embodiments, the reinforcement learning decision layer is used to generate stability-oriented dynamic control strategies in a multi-agent game environment.

[0040] Multi-agent decision-making framework construction: The reinforcement learning decision layer adopts a multi-agent decision-making architecture that combines a centralized critic and distributed actors.

[0041] The joint strategy is defined as: ;in, This indicates the control strategy of the platform's intelligent agent. This represents the behavioral response strategy of a collective intelligent agent of workers. This represents the demand response strategy of the demand-side collective intelligence agent. This represents the constraint application strategy of the system's constraining agent;

[0042] Centralized commentators are based on the system's global state. The system engages in joint actions with multiple agents to evaluate the long-term stability benefits of the current joint strategy; each agent's actuator executes the strategy based only on its local observable information, thus balancing decision-making efficiency and system consistency.

[0043] Please refer to Figure 3. Under the above decision-making framework, Figure 3 shows the joint decision-making structure of each agent. The reinforcement learning decision layer optimizes the joint strategy under the unified stability assessment feedback constraint. The stability assessment result is directly used as the evaluation basis of the joint strategy, which is different from the conventional multi-agent decision-making method with the goal of gain or efficiency.

[0044] Joint policy optimization objective construction: The reinforcement learning decision layer determines the optimal joint control strategy that can maintain the overall stability of the system within the joint policy space. This invention constructs a joint policy optimization objective oriented towards system-level stability, which is formally expressed as follows: ;in, This represents the optimal joint control strategy obtained through reinforcement learning. Let represent the joint policy space consisting of the policies of the platform, the worker group, the demand side, and the system-constrained intelligent agents; t represents the discrete time step of the system operation; and T represents the length of the time window for policy evaluation. This represents the time discount factor, used to balance short-term regulatory effects with long-term system stability. This represents the system-level instantaneous feedback signal output by the stability assessment layer at time step t.

[0045] Through the aforementioned optimization objectives, the reinforcement learning decision layer can guide the system policy to gradually converge to a stable operating range under conditions of multi-party game and dynamic feedback. Unlike existing reinforcement learning methods that focus on returns or task completion rates, the optimization objective of reinforcement learning in this invention is uniformly defined by the stability evaluation layer, ensuring that the policy search direction always revolves around the system's stable operating range.

[0046] Policy update mechanism: In each decision cycle, the reinforcement learning decision layer performs the following process: receiving the system state output by the multi-source data perception layer. Each agent outputs a joint action based on its current policy. ; Receive feedback signals returned by the stability evaluation layer The policy parameters of each agent are updated based on a centralized commentator.

[0047] The process supports online updates and incremental learning to adapt to the environmental nonstationarity and evolutionary characteristics of group behavior in digital workforce systems.

[0048] In some embodiments, the stability assessment layer is used to quantitatively evaluate the operational results of the digital workforce system under a given control strategy, serving as a key intermediary layer connecting strategy learning and closed-loop control. This layer is independent of the reinforcement learning decision-making process and is responsible for calculating stability indicators, identifying risks, and generating feedback signals. The stability assessment layer not only analyzes and evaluates the system's operational status but also directly participates in the system control process as a key intermediary module for reinforcement learning strategy updates and control action constraints.

[0049] Please refer to Figure 4. Figure 4 shows the overall structure of the stability assessment layer and the feedback signal generation process. The stability assessment layer generates system-level feedback signals for updating reinforcement learning strategies through multi-dimensional stability index calculation and comprehensive mapping, and simultaneously applies them to the parameter constraints of the regulation execution stage.

[0050] Stability index system construction: Construct a stability evaluation vector comprising three dimensions: supply side, demand side, and system, represented as follows: ;in, As an indicator of the stability of the working population, As an indicator of demand-side operational stability, The overall coordination and stability index of the system; the stability index of the labor force is defined as: ;in, For the level of worker retention, For the average income of workers, For income volatility, As for the intensity of group attrition or confrontational behavior, , , , Let be the weighting coefficient, satisfying The demand-side operational stability index is defined as follows: ;in, Demand satisfaction level; The intensity of demand fluctuations; , Let be the weighting coefficient, satisfying The system coordination stability index is defined as follows: ;in, For platform control consistency indicators; This is a system constraint satisfaction index; As an indicator of the intensity of multi-party strategic conflicts; , , Let be the weighting coefficient, satisfying ;according to The system's operating state is divided into a stable zone, a transition zone, and a risk zone. When the system enters the risk zone, a risk identification signal is generated. ;in: ; ; ; ; ;in, , , These represent the stable zone, transition zone, and risk zone, respectively. , These represent the upper threshold of the risk zone and the lower threshold of the stability zone, respectively. , These represent the threshold coefficients for the stable region and the threshold coefficients for the risk region, respectively. and These represent the mean and standard deviation of the window stability index, respectively; the stability evaluation results are then processed through a monotonic mapping function. This is transformed into immediate feedback signals for reinforcement learning, used to guide the evolution of the control strategy towards stable system operation, and serves as the basis for policy updates and control parameter constraints, expressed as: ;in, As a comprehensive evaluation result of system stability, It is a monotonic mapping function, which can be a linear function, a piecewise function, or a nonlinear function, used to map a multidimensional stability index to a function of reinforcement learning rewards.

[0051] The feedback signal is used not only for updating strategy parameters but also for dynamically constraining the intensity and action boundaries of regulation during the execution phase. Through this stability reward design, the system can suppress drastic fluctuations in the regulation strategy under short-term disturbances, reduce the oscillation amplitude of the system state between adjacent time steps, and thus improve the stability of system operation.

[0052] In some embodiments, the closed-loop control feedback layer is used to apply the control actions output by the decision layer to the actual platform operating system.

[0053] Strategy-parameter mapping mechanism: The regulation strategy output by reinforcement learning is first mapped to executable parameters on the platform side, including but not limited to incentive intensity adjustment coefficient, order priority weight, service rule tightness parameter, and regulation rhythm and frequency parameter.

[0054] The mapping process follows preset parameter boundaries and compliance constraints to avoid policy outputs causing severe system oscillations.

[0055] Regulation Implementation and Safety Constraints: Before regulation is implemented, the closed-loop feedback layer performs safety checks on the regulation parameters: whether they exceed the platform's acceptable cost range, whether they violate regulatory or compliance constraints, and whether they may cause a drastic imbalance between supply and demand.

[0056] Only when the verification passes will the control action be sent to the actual platform operating system.

[0057] Execution effect monitoring and feedback: The system operation results after the control is executed are collected in real time and fed back to the multi-source data perception layer to update the system status. This feedback information is also used to: verify the effectiveness of the strategy execution; correct model errors; and support subsequent iterative learning of the strategy.

[0058] Through the above mechanism, the system forms a complete closed loop of "strategy generation - regulation execution - status feedback - strategy correction". Closed-loop control is formed at both the operational and strategic levels, thereby avoiding the problem of lag or failure of regulation in a single closed loop under complex game environment, and realizing continuous adaptive regulation of the stability of the digital labor force system.

[0059] Please refer to Figure 5. Figure 5 illustrates the dual-closed-loop control mechanism comprised of the system operation layer and the strategy learning layer. This invention achieves real-time control and long-term stability optimization of the digital workforce system through the synergistic effect of the operation layer closed loop and the strategy learning closed loop. The operation closed loop (right side) enables real-time state control at the second or minute level; the strategy learning closed loop (left side) enables strategy evolution optimization at the hour or day level. The two are cross-coupled at three levels—optimization objective, action constraint, and execution feedback—through stability assessment results, jointly achieving continuous adaptive stability of the system.

[0060] The aforementioned layers are not set up independently and in parallel, but rather form a cross-coupled relationship across the three levels of strategy optimization objectives, regulatory action constraints, and execution feedback through stability assessment results, thereby achieving coordinated regulation of system stability.

[0061] Example 2: Please refer to Figure 1. The closed-loop control system of the digital workforce system based on multi-agent reinforcement learning provided in this application is specifically applied to the ride-hailing service platform scenario. It is used to intelligently control the dynamic supply and demand relationship between the driver group and passenger demand in order to achieve synergistic optimization of platform operation stability and service quality.

[0062] The closed-loop control system of the ride-hailing service platform includes a multi-source data sensing layer for real-time collection of multi-source data from the platform. This multi-source data includes: driver group data (labor supply-side data) to characterize the operational status of ride-hailing drivers, including at least driver online rate (the ratio of online drivers to registered drivers per unit time); order acceptance rate and order rejection rate; order completion rate and cancellation rate; average driver income per order, average daily income, and income fluctuation; and data reflecting driver retention, such as the proportion of drivers continuously offline and the proportion of drivers with periodic churn. Passenger demand-side data (demand-side data) is used to characterize passenger travel patterns. Demand behavior data includes at least the number of orders generated per unit time; demand density distribution in different time periods and regions; passenger sensitivity to price changes and waiting times; records of sudden increases in demand during peak periods, severe weather, or special events; platform control data, including at least order priority rule parameters; dynamic surcharge or subsidy adjustment coefficients; the intensity, frequency, and duration of incentive activities; upper and lower limits of price fluctuations and commission rates; system constraints and historical feedback data, including at least regulatory requirements for compliance with pricing, service quality, etc.; records of historical supply-demand imbalances, driver mass offline events, or service complaint incidents; and historical platform stability assessment results.

[0063] Specifically, the aforementioned multi-source data, after undergoing denoising, normalization, and time alignment processing, is mapped into a unified system state vector. As the input for subsequent multi-agent game modeling and reinforcement learning decision-making, it is represented as: ;in, Indicates the characteristics of the working group's status. This indicates the characteristics of the demand-side group's state; Indicates the characteristics of the platform's control status; This represents the system constraints and historical feedback characteristics.

[0064] The closed-loop control system of the ride-hailing service platform also includes a multi-agent modeling layer, used to model the key participants in the system as four group-level agents; the four group-level agents include: the platform agent. Used to output platform-side control actions, such as dynamic pricing coefficients, driver subsidy intensity, and order dispatch priority parameters; driver swarm intelligence agent. That is, the collective intelligence of the labor force Used to characterize the overall behavioral response of drivers under changes in prices, subsidies, and order dispatch rules, including order acceptance probability and online time trends; passenger group intelligent agents. That is, demand-side collective intelligence Used to describe changes in passengers' willingness to place orders and their cancellation behavior in response to changes in price or waiting time; system constraints on intelligent agents. Used to introduce constraints such as price compliance and service stability to dynamically limit the space for joint strategies.

[0065] Specifically, the key participants in a ride-hailing service platform are defined as a set of four group-level intelligent agents, represented as follows: ;in, It serves as the platform's intelligent agent, used to output platform-side control actions, including dynamic subsidy coefficients, order assignment weights, or rule tightness parameters. It is a collective intelligent agent for drivers, used to characterize the behavioral response tendencies of drivers as a whole or in subgroups, including the probability of accepting orders or the trend of changes in activity levels. For passenger group intelligence agents, it is used to describe the demand elasticity and behavioral feedback of users or employers under changes in price or service quality; For system-constrained agents, it is used to introduce stability, security or compliance constraints that must be met during system regulation. By influencing the system state transition and the joint policy feasible domain, the constraints are embedded in the multi-agent game process.

[0066] The system state transition is described by the following multi-agent Markov game model, and the state transition function is defined as: Where t represents the discrete decision time step, This represents the global state of the system at time step t, output by the multi-source data sensing layer. This indicates the system state at the next moment after the control action is performed. Indicates platform intelligent agent The regulatory actions taken at time step t Represents the driver collective intelligence agent The action taken at time step t. Represents the intelligent agent of the passenger group The demand response actions taken at time step t Represents the system's constrained intelligent agent The constraint action or constraint condition applied at time step t This represents the system state transition probability function, used to describe the evolution of the system state under the joint actions of multiple agents.

[0067] Please refer to Figures 1 and 3. The closed-loop control system of the ride-hailing service platform also includes a reinforcement learning decision layer, which is used to generate a joint control strategy executable by the platform based on multi-agent game modeling. A multi-agent reinforcement learning architecture combining a centralized commentator and distributed executors is adopted.

[0068] In this system, the centralized commentator has access to the global state vector of the system and the joint actions of each agent during the training phase, which is used to evaluate the long-term impact of the current joint strategy on the system stability within a given time window. The distributed actuators are deployed on the platform agent, the driver group agent, and the passenger group agent respectively. During the actual operation phase, they output actions independently based only on their own observable local information, thereby realizing distributed decision execution.

[0069] Specifically, the reinforcement learning decision layer uses the system state vector output by the multi-source data perception layer. For input, It should include at least: online rate, order acceptance rate, and retention rate of the driver group; passenger demand intensity and regional distribution characteristics; current platform price and subsidy parameters; and compliance restrictions on the output of the system's intelligent agent.

[0070] Specifically, based on the system state vector A joint control strategy is generated, wherein the joint control strategy, with the system stability assessment result as the optimization objective, includes: the joint control strategy is defined as a set of strategies of each agent, expressed as: ;in, This indicates the control strategy of the platform's intelligent agent. This represents the behavioral response strategy of the driver collective intelligence agent. This represents the demand response strategy of the passenger group's intelligent agent. This represents the constraint-imposing strategy of the system's constrained agents; centralized commentators are based on... By coordinating actions with multiple agents, the long-term stability benefits of the current joint strategy are evaluated. Each agent's actuator outputs the corresponding action based only on its local observable information, achieving distributed decision-making and execution. The optimization objective of the joint control strategy is constructed as follows: ;in, This represents the optimal joint control strategy obtained through reinforcement learning. Let represent the joint policy space consisting of the policies of the platform, driver group, passengers, and system-constrained agents, where t represents the discrete time step of system operation, and T represents the length of the time window for policy evaluation. This represents the time discount factor, used to balance short-term regulatory effects with long-term system stability. This represents the system-level instantaneous feedback signal output by the stability assessment layer at time step t.

[0071] The joint strategy output should include at least: platform-side control actions, such as dynamic pricing adjustment range, driver subsidy intensity range, and order priority weight parameters; driver group behavior response prediction results, such as changes in overall order acceptance tendency and online time trend; passenger group demand response prediction results, such as changes in order probability and cancellation rate trend; and system constraint conditions on the feasibility of the above actions.

[0072] The optimization objective of the joint control strategy is to maximize system stability. Specifically, within a given time window, through iterative reinforcement learning strategies, the system aims to maintain stable performance in terms of supply and demand balance, driver retention, service quality, and compliance within a predetermined range. By introducing a time discount factor, the reinforcement learning decision-making layer can balance short-term control effects with long-term system stability, avoiding system fluctuations caused by excessive short-term optimization.

[0073] Please refer to Figure 4. The closed-loop control system of the ride-hailing service platform also includes a stability assessment layer, which is used to conduct multi-dimensional stability assessment of the operation status of the ride-hailing service platform and generate feedback signals required for reinforcement learning.

[0074] The stability assessment layer constructs a system stability index system from the following three dimensions: driver group stability index: used to reflect the overall stability of the driver supply side, the index includes at least the driver retention rate and its trend; the average income level of drivers and the magnitude of income fluctuation; and the intensity of abnormal behaviors such as group order rejection and collective offline.

[0075] Passenger demand-side operational stability indicators: These indicators reflect the smoothness of demand-side operations and include at least the order fulfillment rate; the degree of fluctuation in waiting time; and the intensity of demand fluctuations in different time periods and regions.

[0076] System overall coordination and stability indicators: These indicators are used to characterize the coordination and consistency of platform control strategies among multiple parties. The indicators include at least the consistency and continuity of platform control actions; the degree to which regulatory constraints and compliance conditions are met; and the intensity of conflict between platform control behavior and driver / passenger behavior.

[0077] Specifically, the system stability level is calculated based on driver group stability indicators, passenger demand-side operational stability indicators, and overall system coordination stability indicators, and then mapped to reinforcement learning real-time feedback signals. This includes: constructing a stability evaluation vector encompassing three dimensions: driver supply, passenger demand, and the system itself, represented as: ;in, As an indicator of driver group stability, For passenger demand-side operational stability indicators, The overall system coordination stability index; the driver group stability index is defined as: ;in, For driver retention level, For the average income of drivers, For income volatility, As for the intensity of group attrition or confrontational behavior, , , , Let be the weighting coefficient, satisfying The passenger demand-side operational stability index is defined as: ;in, Demand satisfaction level; The intensity of demand fluctuations; , Let be the weighting coefficient, satisfying The system coordination stability index is defined as follows: ;in, For platform control consistency indicators; This is a system constraint satisfaction index; As an indicator of the intensity of multi-party strategic conflicts; , , Let be the weighting coefficient, satisfying .

[0078] The stability assessment layer integrates the indicators from the three dimensions mentioned above to calculate the current stability level of the system and divides the system state into a stable zone, a transition zone, and a risk zone. When the system state enters the risk zone, a risk indicator signal is generated to prompt the reinforcement learning decision layer to update the control strategy to a more conservative or more restrictive one. Finally, the stability assessment result is converted into an immediate feedback signal for reinforcement learning through a pre-defined mapping function, serving as the core basis for optimizing the joint control strategy.

[0079] according to The system's operating state is divided into a stable zone, a transition zone, and a risk zone; when the system enters the risk zone, a risk identification signal is generated. ;in: ; ; ; ; ;in, , , These represent the stable zone, transition zone, and risk zone, respectively. , These represent the upper threshold of the risk zone and the lower threshold of the stability zone, respectively. , These represent the threshold coefficients for the stable region and the threshold coefficients for the risk region, respectively. and These represent the mean and standard deviation of the window stability index, respectively; the stability evaluation results are then processed through a monotonic mapping function. This is transformed into immediate feedback signals for reinforcement learning, used to guide the evolution of the control strategy towards stable system operation, and serves as the basis for policy updates and control parameter constraints, expressed as: ;in, As a comprehensive evaluation result of system stability, It is a monotonic mapping function, which can be a linear function, a piecewise function, or a nonlinear function, used to map a multidimensional stability index to a function of reinforcement learning rewards.

[0080] Please refer to Figure 5. The closed-loop control system of the ride-hailing service platform also includes a closed-loop control feedback layer, which is used to map the control strategy output by reinforcement learning into control actions that the platform can execute. Before execution, a safety verification is performed, and the system operation results after control are fed back to the multi-source data perception layer in real time, forming a closed-loop control mechanism.

[0081] The closed-loop control feedback layer includes the following steps: Strategy mapping: The joint control strategy output by the reinforcement learning decision layer is mapped to control parameters that the ride-hailing service platform can directly execute. The control parameters include, but are not limited to: dynamic price increase or subsidy adjustment coefficient; order dispatch priority and matching weight; and the rhythm and duration of incentive activities.

[0082] Security verification: Before the control parameters are issued, the parameters are verified for security. The verification content includes at least: whether they exceed the operating cost range that the platform can bear; whether they violate regulatory or compliance constraints such as pricing and service quality; and whether they may cause a drastic imbalance between supply and demand or abnormal group behavior.

[0083] The control parameters are only sent to the actual operating system for execution when the security check passes.

[0084] Feedback flow: The system operation results after the control is executed are collected in real time and re-input to the multi-source data perception layer to construct the system state vector at the next moment, thereby driving subsequent policy evaluation and updates, forming a complete closed-loop control mechanism.

[0085] The aforementioned layers are not set up independently and in parallel, but rather form a cross-coupled relationship across the three levels of strategy optimization objectives, regulatory action constraints, and execution feedback through stability assessment results, thereby achieving coordinated regulation of system stability.

[0086] Example 3: Based on the same inventive concept, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the closed-loop control method for a digital workforce system based on multi-agent reinforcement learning as described above.

[0087] The program product of this application for implementing the above method may employ a portable compact disk read-only memory and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0088] It should be noted that a computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0089] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although this application has disclosed preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-mentioned technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. The implementation schemes in the above embodiments can also be further combined or replaced. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the content of the technical solution of this application shall still fall within the scope of this application.

Claims

1. A closed-loop control method for a digital workforce system based on multi-agent reinforcement learning, characterized in that, include: S1. Multi-source data perception and state construction: Real-time acquisition of multi-source data from the digital workforce system, preprocessing and fusion, and construction of a unified system state vector. The multi-source data includes at least labor supply-side data, demand-side data, platform regulation data, and system constraint and historical feedback data; S2, Multi-agent game modeling: The key participants in the digital labor system are modeled as four group-level agents, namely, platform agents. Labor group intelligence Demand-side collective intelligence and system-constrained intelligent agents Furthermore, a multi-party Markov game model is constructed to describe the dynamic transition of system states and the interaction of multi-party strategies; S3, Reinforcement Learning Decision Generation: A multi-agent reinforcement learning architecture combining a centralized commentator and a distributed executor is adopted, based on the system state vector. A joint control strategy is generated, with the system stability assessment result as the optimization objective; S4, Stability assessment and feedback signal generation: Based on the stability index of the worker group, the demand-side operation stability index, and the overall system coordination stability index, the system stability level is calculated and mapped to reinforcement learning real-time feedback signals. ; S5. Closed-loop control execution and feedback: The joint control strategy output by reinforcement learning is mapped to platform executable parameters. Security verification is performed before execution, and the system operation results after control are fed back to multi-source data perception in real time, forming a closed-loop control mechanism from S1 to S5.

2. The method according to claim 1, characterized in that, The multi-source data acquired in the real-time digital workforce system is preprocessed and fused to construct a unified system state vector. The data includes: labor supply-side data such as worker order acceptance rate, online rate, task completion rate, average income and fluctuation range, order rejection or abnormal behavior ratio, and historical retention and churn trends; demand-side data such as order generation intensity and time distribution, user sensitivity to price and waiting time, demand elasticity and fluctuation characteristics, and historical demand stability records; platform control data such as order dispatch rule parameters, subsidy or incentive intensity range, and rule adjustment frequency and magnitude; and system constraint and historical feedback data such as historical system stability assessment results, compliance constraint information, and risk event records. After denoising, normalizing, and time-aligning, the multi-source data is mapped to a unified system state vector. , is represented as: ;in, Indicates the characteristics of the working group's status. This indicates the characteristics of the demand-side group's state; Indicates the characteristics of the platform's control status; This represents the system constraints and historical feedback characteristics.

3. The method according to claim 2, characterized in that, The key stakeholders in the digital workforce system are modeled as four group-level intelligent agents, namely the platform intelligent agent. Labor group intelligence Demand-side collective intelligence and system-constrained intelligent agents Furthermore, a multi-party Markov game model is constructed to describe the dynamic transition of system states and the interaction of multi-party strategies. This includes defining the key participants in the digital workforce system as a set of four group-level agents, denoted as: ;in, It serves as the platform's intelligent agent, used to output platform-side control actions, including dynamic subsidy coefficients, order assignment weights, or rule tightness parameters. It is a collective intelligence agent for workers, used to characterize the behavioral response tendencies of workers as a whole or in subgroups, including the probability of order acceptance or the trend of activity changes. It is a demand-side collective intelligence agent used to describe the demand elasticity and behavioral feedback of users or employers under changes in price or service quality. For system-constrained agents, constraints are introduced to define stability, security, or compliance requirements that must be met during system regulation. These constraints are embedded into the multi-agent game process by influencing system state transitions and the feasible region of the joint strategy. The system state transitions are described by the following multi-agent Markov game model, and the state transition function is defined as: Where t represents the discrete decision time step, This represents the global state of the system at time step t, output by the multi-source data sensing layer. This indicates the system state at the next moment after the control action is performed. Indicates the platform's intelligent agent The regulatory actions taken at time step t Represents the collective intelligence of the working class The action taken at time step t. Represents the demand-side collective intelligence agent The demand response actions taken at time step t Represents the system's constrained intelligent agent The constraint action or constraint condition applied at time step t This represents the system state transition probability function, used to describe the evolution of the system state under the joint actions of multiple agents.

4. The method according to claim 2, characterized in that, The multi-agent reinforcement learning architecture, which combines a centralized commentator with a distributed executor, is based on the system state vector. A joint control strategy is generated, wherein the joint control strategy, with the system stability assessment result as the optimization objective, includes: the joint control strategy is defined as a set of strategies of each agent, expressed as: ;in, This indicates the control strategy of the platform's intelligent agent. This represents the behavioral response strategy of a collective intelligent agent of workers. This represents the demand response strategy of the demand-side collective intelligence agent. This represents the constraint-imposing strategy of the system's constrained agents; centralized commentators are based on... By coordinating actions with multiple agents, the long-term stability benefits of the current joint strategy are evaluated. Each agent's actuator outputs the corresponding action based only on its local observable information, achieving distributed decision-making and execution. The optimization objective of the joint control strategy is constructed as follows: ;in, This represents the optimal joint control strategy obtained through reinforcement learning. Let represent the joint policy space consisting of the policies of the platform, the worker group, the demand side, and the system-constrained intelligent agents; t represents the discrete time step of the system operation; and T represents the length of the time window for policy evaluation. This represents the time discount factor, used to balance short-term regulatory effects with long-term system stability. This represents the system-level instantaneous feedback signal output by the stability assessment layer at time step t.

5. The method according to claim 4, characterized in that, The system stability level is calculated based on the stability index of the labor force, the stability index of demand-side operation, and the overall coordination stability index of the system, and then mapped to reinforcement learning real-time feedback signals. This includes: constructing a stability evaluation vector encompassing three dimensions: supply side, demand side, and system, represented as: ;in, As an indicator of the stability of the working population, As an indicator of demand-side operational stability, The overall coordination and stability index of the system; the stability index of the labor force is defined as: ;in, For the level of worker retention, For the average income of workers, For income volatility, As for the intensity of group attrition or confrontational behavior, , , , Let be the weighting coefficient, satisfying The demand-side operational stability index is defined as follows: ;in, Demand satisfaction level; The intensity of demand fluctuations; , Let be the weighting coefficient, satisfying The system coordination stability index is defined as follows: ;in, For platform control consistency indicators; This is a system constraint satisfaction index; As an indicator of the intensity of multi-party strategic conflicts; , , Let be the weighting coefficient, satisfying ;according to The system's operating state is divided into a stable zone, a transition zone, and a risk zone. When the system enters the risk zone, a risk identification signal is generated. ;in: ; ; ; ; ;in, 、 、 These represent the stable zone, transition zone, and risk zone, respectively. 、 These represent the upper threshold of the risk zone and the lower threshold of the stability zone, respectively. 、 These represent the threshold coefficients for the stable region and the threshold coefficients for the risk region, respectively. and These represent the mean and standard deviation of the window stability index, respectively; the stability evaluation results are then processed through a monotonic mapping function. This is transformed into immediate feedback signals for reinforcement learning, used to guide the evolution of the control strategy towards stable system operation, and serves as the basis for policy updates and control parameter constraints, expressed as: ;in, As a comprehensive evaluation result of system stability, It is a monotonic mapping function, which can be a linear function, a piecewise function, or a nonlinear function, used to map a multidimensional stability index to a function of reinforcement learning rewards.

6. The method according to claim 5, characterized in that, The mechanism of mapping the joint control strategy output by reinforcement learning into executable parameters on the platform, performing security checks before execution, and feeding back the system operation results after control to multi-source data perception in real time to form a closed-loop control mechanism includes: mapping the joint control strategy output by reinforcement learning decision into executable control parameters on the platform side, wherein the control parameters include incentive intensity adjustment coefficient, order priority weight, service rule tightness parameter, and control rhythm and frequency parameter; performing security checks before the control parameters are issued to the actual operating system, wherein the security checks include at least whether it exceeds the platform's acceptable cost range, whether it violates regulatory or compliance constraints, and whether it may cause a drastic imbalance between supply and demand; the control action is issued and executed only when the security checks pass; the system operation results after control execution are collected in real time and fed back to multi-source data perception to construct the system state at the next moment. This information is then used for subsequent strategy iterations and learning.

7. A closed-loop control system for a digital workforce system based on multi-agent reinforcement learning, characterized in that: The digital workforce system stability control system includes: a multi-source data sensing layer, used to collect multi-source data from the digital workforce system in real time, perform preprocessing and fusion, and construct a unified system state vector. The multi-source data includes at least labor supply-side data, demand-side data, platform regulation data, and system constraint and historical feedback data. The multi-agent modeling layer is used to model the key participants in the digital workforce system as four group-level agents, namely, platform agents. Labor group intelligence Demand-side collective intelligence and system-constrained intelligent agents A multi-party Markov game model is constructed to describe the dynamic transition of system states and the interaction of multi-party strategies; a reinforcement learning decision layer is used to employ a multi-agent reinforcement learning architecture that combines a centralized commentator and distributed executors, based on the system state vector. A joint control strategy is generated, with the system stability assessment result as the optimization objective. A stability assessment layer is used to calculate the system stability level based on worker group stability indicators, demand-side operational stability indicators, and overall system coordination stability indicators, and maps this level to reinforcement learning real-time feedback signals. The closed-loop control feedback layer is used to map the joint control strategy output by reinforcement learning into platform executable parameters, perform security verification before execution, and feed back the system operation results after control to the multi-source data perception in real time, forming a closed-loop control mechanism from the multi-source data perception layer to the closed-loop control execution and feedback layer.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it is used to implement the closed-loop control method for a digital labor force system based on multi-agent reinforcement learning as described in any one of claims 1-6.