Team cooperation training method and system based on scene reproduction and multi-agent cooperation
By constructing multi-agent interactive scenarios and deep reinforcement learning, combining knowledge distillation and counterfactual reasoning, and dynamically adjusting the training strategy, the gap between scenarios and actual needs and inaccurate evaluation in team collaboration training is solved, and efficient collaboration skills training is achieved.
Patent Information
- Application Number
- CN202510857525.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing team collaboration training system lacks effective utilization of real historical data, the training scenarios are far from the actual needs, and the precise evaluation mechanism and adaptive learning ability are lacking, resulting in poor training results.
By establishing a team collaboration knowledge base, building a multi-agent interaction scenario, using deep reinforcement learning to train the behavioral patterns of agents, combining the multi-agent collaborative learning network enhanced by knowledge distillation, conducting counterfactual reasoning analysis, dynamically quantifying the impact of cooperative behavior on task results, generating a collaboration evaluation report and adjusting training strategies.
It has achieved the authenticity and targeted improvement of team collaboration training, dynamically adjust the training difficulty, accurately evaluate the collaboration effect, shorten the running-in period, and improve collaboration efficiency.
Smart Images

Figure CN120374058A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to team collaboration technology, and particularly to a team collaboration training method and system based on scenario reproduction and multi-agent collaboration. Background Art
[0002] The ability of team collaboration is a crucial core competitiveness in modern organizations. Especially in complex task environments, effective team collaboration can significantly improve work efficiency and the quality of task completion. Traditional team collaboration training mainly relies on methods such as case analysis, role-playing, and on-site drills. Although these methods can, to a certain extent, enhance the collaboration awareness of team members, they often lack pertinence and precision. With the rapid development of artificial intelligence technology, team collaboration training methods based on multi-agent systems have gradually emerged. This training method can simulate real task scenarios and provide a richer and more accurate training experience.
[0003] In the field of team collaboration training, existing technologies usually adopt static scenario design and simplified interaction modes, which are difficult to truly reflect the collaboration challenges of teams in complex environments. At the same time, traditional training methods generally have problems such as imperfect effect evaluation systems, disconnection between training scenarios and actual work, and lack of personalized adaptation in the training process, which limit the effective improvement of team collaboration ability.
[0004] Existing team collaboration training systems lack the effective utilization of real historical data. Training scenarios are often preset standardized environments and cannot accurately reproduce the specific challenge situations encountered by teams in actual work, resulting in a large gap between training content and actual needs, and it is difficult to transform training effects into actual work capabilities.
[0005] Traditional collaboration training methods lack an accurate quantitative evaluation mechanism and cannot objectively measure the synergy between team members and its impact on task results. Training feedback usually relies on subjective judgment, making it difficult to identify key behaviors and ability shortboards in the collaboration process, resulting in the lack of precision and pertinence in training adjustment.
[0006] Existing team collaboration training systems generally lack the ability of adaptive learning, cannot dynamically adjust training content and difficulty according to the performance and ability levels of team members, and cannot capture the evolution trajectory of team capabilities, resulting in low training efficiency and difficulty in meeting the needs of different stages of team development and the cultivation of differentiated collaboration capabilities. Summary of the Invention
[0007] Embodiments of the present invention provide a team collaboration training method and system based on scenario reproduction and multi-agent collaboration, which can solve the problems in the prior art.
[0008] In the first aspect of the embodiments of the present invention, a team collaboration training method based on scenario reproduction and multi-agent collaboration is provided, including: Build a team collaboration knowledge base based on pre - obtained team collaboration historical data, construct a multi - agent interaction scenario according to the team collaboration knowledge base, map the virtual environment parameters, agent role information, and task objective information in the team collaboration knowledge base to the multi - agent interaction scenario, and generate a scenario reproduction training environment; Generate agents in the scenario reproduction training environment, and train the initial behavior patterns of the agents through a deep reinforcement learning model based on the historical behavior data in the team collaboration knowledge base; Construct a multi - agent collaborative learning network enhanced by knowledge distillation, and feed the collaborative behavior data generated by team members interacting with the agents in the scenario reproduction training environment to the multi - agent collaborative learning network in real - time; Analyze the influence degree of team collaboration behavior on task results through a counterfactual reasoning method based on the collaborative behavior data, construct a causal relationship graph between collaborative behavior and task performance, dynamically quantify the synergy strength among team members in the causal relationship graph, and combine the temporal attention mechanism to track the ability evolution trajectory; Generate a team collaboration evaluation report according to the ability evolution trajectory; based on the team collaboration evaluation report, output a team collaboration optimization plan, adjust the parameter configuration of the scenario reproduction training environment specifically, and update the training strategy of the multi - agent collaborative learning network.
[0009] Mapping the virtual environment parameters, agent role information, and task objective information in the team collaboration historical data to the multi - agent interaction scenario to generate a scenario reproduction training environment includes: Construct a feature vector space mapping model, extract features from the virtual environment parameters, agent role information, and task objective information based on the feature vector space mapping model to generate a target scenario feature vector, and the target scenario feature vector includes an environment feature component, a role feature component, and a task feature component; Construct a multi - objective optimization function based on the target scenario feature vector, calculate the feature mapping loss for the environment feature component, the role feature component, and the task feature component respectively according to the multi - objective optimization function, and optimize the target scenario feature vector based on the weighted result of the feature mapping loss; Input the optimized target scenario feature vector into the generator of the generative adversarial network, and the generator reconstructs the features of the target scenario feature vector through a multi - layer transposed convolutional network to generate an initial scenario structure including environment layout parameters, role attribute parameters, and task configuration parameters; Construct a scene difficulty adaptive adjustment module based on the initial scene structure. The scene difficulty adaptive adjustment module calculates a scene difficulty coefficient according to a preset performance index, and dynamically adjusts the initial scene structure based on the scene difficulty coefficient; Instantiate the adjusted initial scene structure into a training environment, and configure an agent interaction interface and a task evaluation module in the training environment. The task evaluation module dynamically adjusts the evaluation criteria based on the scene difficulty coefficient to achieve the adaptive generation of the scene reproduction training environment.
[0010] Construct a multi-objective optimization function based on the target scene feature vector, and calculate the feature mapping losses for the environmental feature component, the role feature component, and the task feature component respectively according to the multi-objective optimization function, including: Construct a feature component representation model, and encode the environmental feature component, the role feature component, and the task feature component into feature vectors respectively based on the feature component representation model to generate an initial feature vector set; Construct a multi-objective mapping loss function based on the initial feature vector set. The multi-objective mapping loss function performs weighted calculation on the environmental feature mapping loss, the role feature mapping loss, and the task feature mapping loss to generate a basic mapping loss value; Establish a feature causal graph structure using the initial feature vector set. The nodes of the feature causal graph structure are the feature vectors in the initial feature vector set, and the edges of the feature causal graph structure represent the causal relationships between the feature vectors. Calculate the causal strength between the feature nodes through causal intervention operations; Fuse the causal strength with the basic mapping loss value to construct a causally enhanced overall loss function; calculate the parameter gradient based on the overall loss function, and update the parameters of the feature component representation model based on the parameter gradient; Calculate the corresponding weight coefficients for the environmental feature mapping loss, the role feature mapping loss, and the task feature mapping loss based on the updated feature component representation model to obtain the feature mapping losses of each feature component.
[0011] Generate an agent in the scene reproduction training environment, and train the initial behavior pattern of the agent through a deep reinforcement learning model based on the historical behavior data in the team collaboration knowledge base, including: Generate an agent in the scene reproduction training environment, extract historical behavior data from the team collaboration knowledge base, and construct a deep reinforcement learning model based on the historical behavior data. The deep reinforcement learning model includes a behavior state space, an action space, and a reward function; Iteratively train the agent using the deep reinforcement learning model, update the policy network parameters of the agent according to the feedback information of the reward function, and train the initial behavior pattern of the agent.
[0012] Based on the collaborative behavior data, analyze the influence degree of team collaborative behavior on task results through the counterfactual reasoning method, and construct a causal relationship graph between collaborative behavior and task performance. The causal relationship graph dynamically quantifies the synergy strength among team members. Combining the time-series attention mechanism to track the ability evolution trajectory includes: Construct a collaborative behavior vector and a task result vector based on the collaborative behavior data, and use the causal intervention operator to calculate the causal effect of each behavior feature in the collaborative behavior vector on each result index in the task result vector. The causal effect characterizes the influence strength of the collaborative behavior on the task result, and construct a causal relationship graph between team collaborative behavior and task performance based on the causal effect; Calculate the synergy strength among team members based on the causal relationship graph, combine the time-series weight and the attention weight to weight the causal effect, and obtain the dynamic synergy strength; Construct a team ability phase space based on the dynamic synergy strength, calculate an ability index based on the team ability phase space. The ability index is used to characterize the stability of team ability evolution, and input the ability index into the multi-scale time-series attention mechanism; The multi-scale time-series attention mechanism calculates the dynamic correlation weight of time-series features using a query matrix, a key matrix, and a value matrix, and constructs a Jacobian matrix based on the dynamic correlation weight. The Jacobian matrix describes the local linear characteristics of the ability phase space; Combine the ability index and the Jacobian matrix to calculate a bifurcation discriminant, predict the transition probability of team ability based on the bifurcation discriminant. The transition probability indicates the turning point of ability evolution; combine the dynamic correlation weight and the transition probability to track and predict the team ability evolution trajectory.
[0013] Combining the ability index and the Jacobian matrix to calculate a bifurcation discriminant, and predicting the transition probability of team ability based on the bifurcation discriminant includes: Combine the ability index and the eigenvalues of the Jacobian matrix to construct a bifurcation parameter vector. The bifurcation parameter vector contains the maximum real part of the ability index and the eigenvalues; Based on the bifurcation parameter vector, calculate the bifurcation discriminant through the determinant of the Jacobian matrix and the exponential smoothing term. The bifurcation discriminant is used to detect the mutation characteristics of the team ability state; Calculate the critical value of the bifurcation discriminant within the observation time window. The critical value characterizes the threshold condition for the transition of the team's ability state, and input the difference between the bifurcation discriminant and the critical value into a non-linear mapping function; The non-linear mapping function is constructed based on the sigmoid function to calculate the transition probability of the team's ability. The transition probability characterizes the possibility of the mutation of the team's ability; Construct the confidence interval of the transition probability. The confidence interval calculates the upper and lower bounds through an uncertainty function, and the uncertainty function is dynamically adjusted with the change of the bifurcation discriminant to achieve accurate prediction of the team's ability transition.
[0014] Generate a team collaboration evaluation report according to the ability evolution trajectory; based on the team collaboration evaluation report, output a team collaboration optimization plan, and specifically adjust the parameter configuration of the scenario reproduction training environment, and update the training strategy of the multi-agent collaborative learning network, including: The ability evolution trajectory includes multi-dimensional ability indicators. Calculate the individual ability growth rate and the team collaboration efficiency based on the multi-dimensional ability indicators. The individual ability growth rate characterizes the ability improvement speed, and the team collaboration efficiency characterizes the collaboration effect; perform a weighted combination of the individual ability growth rate and the team collaboration efficiency to obtain a team collaboration comprehensive score, and the team collaboration comprehensive score reflects the overall performance of the team; Construct a multi-dimensional evaluation system based on the team collaboration comprehensive score. The multi-dimensional evaluation system includes an individual dimension, a team dimension, and a task dimension, and generate a team collaboration evaluation report; construct an optimization objective function according to the team collaboration evaluation report. The optimization objective function combines performance loss and optimization cost, and generates a team collaboration optimization plan based on resource constraints, time constraints, and effect constraints; Convert the team collaboration optimization plan into the parameter configuration of the scenario reproduction training environment. The parameter configuration includes training difficulty parameters and scenario complexity parameters, dynamically adjust the scenario reproduction training environment, and update the training strategy of the multi-agent collaborative learning network based on the parameter configuration.
[0015] In the second aspect of the embodiments of the present invention, a team collaboration training system based on scenario reproduction and multi-agent collaboration is provided, including: A first unit for establishing a team collaboration knowledge base based on pre-acquired team collaboration historical data, constructing a multi-agent interaction scenario according to the team collaboration knowledge base, mapping the virtual environment parameters, agent role information, and task target information in the team collaboration knowledge base to the multi-agent interaction scenario, and generating a scenario reproduction training environment; The second unit is used to generate an agent in the scenario reproduction training environment, and train the initial behavior pattern of the agent through a deep reinforcement learning model based on the historical behavior data in the team collaboration knowledge base; The third unit is used to construct a knowledge distillation-enhanced multi-agent collaborative learning network, and feedback the collaborative behavior data generated by team members interacting with the agent in the scenario reproduction training environment to the multi-agent collaborative learning network in real time; The fourth unit is used to analyze the influence degree of team collaboration behavior on the task result through a counterfactual reasoning method based on the collaborative behavior data, construct a causal relationship diagram between collaborative behavior and task performance, dynamically quantify the collaborative effect strength among team members by the causal relationship diagram, and track the ability evolution trajectory by combining the temporal attention mechanism; The fifth unit is used to generate a team collaboration evaluation report according to the ability evolution trajectory; based on the team collaboration evaluation report, output a team collaboration optimization plan, adjust the parameter configuration of the scenario reproduction training environment pertinently, and update the training strategy of the multi-agent collaborative learning network.
[0016] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the foregoing method.
[0017] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.
[0018] The beneficial effects of this application are as follows: Through the training method of scenario reproduction and multi-agent collaboration, the present invention can accurately simulate the real team collaboration scenario, provide an immersive collaborative skill training experience for team members, greatly improve the authenticity and pertinence of team collaboration training, thereby effectively shortening the team running-in period and improving the collaboration efficiency.
[0019] The present invention adopts a technical architecture combining deep reinforcement learning and knowledge distillation, realizes the continuous optimization of the agent behavior pattern and the adaptive adjustment of the environment parameters, can dynamically adjust the training difficulty and collaboration requirements according to the actual performance of team members, ensure that the training content is always in the "zone of proximal development" of team members, and significantly improve the personalization level and learning efficiency of the training effect.
[0020] The counterfactual reasoning analysis framework and the temporal attention mechanism constructed in the present invention can accurately quantify the contribution degrees of different collaboration behaviors to the task results and the synergy strength among team members, form an objective and detailed team collaboration evaluation report and optimization plan, provide data-driven decision-making support for team building and the improvement of collaboration capabilities, and effectively solve the problems of fuzzy evaluation criteria and unclear optimization directions in traditional team training. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a schematic flowchart of the team collaboration training method based on scenario reproduction and multi-agent collaboration according to an embodiment of the present invention; Figure 2 is a bar chart comparing the task completion rates of the scenario difficulty adaptation method according to an embodiment of the present invention; Figure 3 is a bar chart comparing the multi-objective optimization effects of the feature components according to an embodiment of the present invention; Figure 4 is a logic block diagram of team collaboration behavior analysis and ability evolution tracking according to an embodiment of the present invention; Figure 5 is a bar chart comparing the analysis of the probability of different team ability transitions according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0024] Figure 1 is a schematic flowchart of the team collaboration training method based on scenario reproduction and multi-agent collaboration according to an embodiment of the present invention, as Figure 1 shown, the method includes: Establish a team collaboration knowledge base based on the pre-acquired team collaboration historical data, construct a multi-agent interaction scenario according to the team collaboration knowledge base, and map the virtual environment parameters, agent role information, and task objective information in the team collaboration knowledge base to the multi-agent interaction scenario to generate a scenario reproduction training environment; Generate an intelligent agent in the scenario reproduction training environment, and train the initial behavior pattern of the intelligent agent through a deep reinforcement learning model based on the historical behavior data in the team collaboration knowledge base; Construct a multi-agent collaborative learning network enhanced by knowledge distillation, and feed back the collaborative behavior data generated by the interaction between the team members and the agent in the scene reproduction training environment to the multi-agent collaborative learning network in real time; Based on the collaborative behavior data, the influence of team collaborative behavior on task results is analyzed by counterfactual reasoning method, and a causal relationship diagram between collaborative behavior and task performance is constructed. The causal relationship diagram dynamically quantifies the intensity of synergy between team members, and combines the temporal attention mechanism to track the evolution trajectory of capabilities; According to the capability evolution trajectory, a team collaboration evaluation report is generated; based on the team collaboration evaluation report, a team collaboration optimization plan is output, the parameter configuration of the scenario reproduction training environment is adjusted in a targeted manner, and the training strategy of the multi-agent collaborative learning network is updated.
[0025] In an optional implementation, mapping the virtual environment parameters, agent role information, and task target information in the team collaboration history data to the multi-agent interaction scene to generate a scene reproduction training environment includes: Constructing a feature vector space mapping model, extracting features of virtual environment parameters, agent role information, and task target information based on the feature vector space mapping model, and generating a target scene feature vector, wherein the target scene feature vector includes an environment feature component, a role feature component, and a task feature component; Constructing a multi-objective optimization function based on the target scene feature vector, calculating feature mapping losses for the environment feature component, the role feature component, and the task feature component according to the multi-objective optimization function, and optimizing the target scene feature vector based on a weighted result of the feature mapping losses; The optimized target scene feature vector is input into the generator of the generative adversarial network, and the generator reconstructs the target scene feature vector through a multi-layer transposed convolutional network to generate an initial scene structure including environment layout parameters, role attribute parameters and task configuration parameters; Building a scene difficulty adaptive adjustment module based on the initial scene structure, the scene difficulty adaptive adjustment module calculates a scene difficulty coefficient according to a preset performance indicator, and dynamically adjusts the initial scene structure based on the scene difficulty coefficient; The adjusted initial scene structure is instantiated as a training environment, and an agent interaction interface and a task evaluation module are configured in the training environment. The task evaluation module dynamically adjusts the evaluation criteria based on the scene difficulty coefficient to achieve adaptive generation of a scene reproduction training environment.
[0026] Build a feature vector space mapping model, which adopts a bidirectional long short-term memory network structure and includes an environment encoding layer, a role encoding layer, and a task encoding layer. The environment encoding layer processes virtual environment parameters, such as terrain information, climate conditions, obstacle distribution, etc., using a 1024-dimensional hidden state vector; the role encoding layer represents the intelligent agent's role attributes using a 512-dimensional embedding vector, including ability values, behavior patterns, decision-making preferences, etc.; the task encoding layer processes task objective information through a 768-dimensional task representation vector, including task types, completion conditions, time constraints, etc.
[0027] The model integrates the encoding information of the three layers through an attention mechanism to generate a 2304-dimensional target scene feature vector, where the environmental feature component accounts for 1024 dimensions, the role feature component accounts for 512 dimensions, and the task feature component accounts for 768 dimensions.
[0028] Based on the generated target scene feature vector, the system constructs a multi-objective optimization function to optimize the feature mapping. This optimization function considers three objectives: environmental consistency, role behavior rationality, and task challenge. The environmental consistency objective is evaluated by calculating the structural similarity between the generated environment and the historical environment, and convolutional operations are used to extract the environmental topological features; the role behavior rationality objective is quantified by comparing the deviation degree between the generated role behavior pattern and the historical intelligent agent's behavior trajectory; the task challenge objective evaluates the matching degree between the task difficulty and the historical task based on the completion rate and average time used.
[0029] The feature mapping losses generated by each objective calculation are integrated through adaptive weights. The weight of the environmental feature loss is 0.4, the weight of the role feature loss is 0.35, and the weight of the task feature loss is 0.25. The optimization process uses the gradient descent method to iterate 5000 times, the learning rate is set to 0.001, and it decays by 10% every 500 iterations until the comprehensive loss of the target scene feature vector drops below the preset threshold of 0.05.
[0030] The optimized target scene feature vector is input into the generator of the generative adversarial network for feature reconstruction. The generator adopts a five-layer transposed convolutional network architecture. The first layer receives the 2304-dimensional target scene feature vector and expands the features through 128 5×5 convolutional kernels; the second layer uses 64 4×4 transposed convolutional kernels for upsampling; the third layer applies 32 3×3 convolutional kernels to generate the basic contour of the scene; the fourth layer uses 16 3×3 convolutional kernels to refine the scene details; the fifth layer uses 8 2×2 convolutional kernels to generate the final scene layout.
[0031] After each layer of transposed convolution, batch normalization and the LeakyReLU activation function are connected, and the activation parameter is set to 0.2. The generator output includes an environmental layout parameter matrix with a dimension of 128×128×3, representing terrain height, material type, and obstacle density in three-dimensional space; a character attribute parameter matrix with a dimension of 64×64×6, representing the position, orientation, speed, energy value, communication range, and skill cooldown time of the agent; and a task configuration parameter matrix with a dimension of 32×32×4, representing task trigger points, target areas, resource distributions, and time nodes.
[0032] Based on the generated initial scene structure, a scene difficulty adaptive adjustment module is implemented. This module uses reinforcement learning methods to dynamically adjust scene parameters to adapt to different user ability levels. The module first defines a scene difficulty coefficient, with an initial value set to 0.5 (medium difficulty) and a range of 0.1 - 0.9. The system calculates the adaptive difficulty based on three preset performance metrics: task completion rate (target value 85%), average decision-making time (target value 1.5 seconds), and collaboration efficiency index (target value 0.7).
[0033] When the user's performance exceeds the target value, the difficulty coefficient increases by 0.05 per round; otherwise, it decreases by 0.05. The scene adjustment follows specific rules: when the difficulty coefficient increases, the obstacle density in the environmental layout parameters increases by 20%, and the terrain complexity increases by 15%; the energy value in the character attribute parameters decreases by 10%, and the skill cooldown time is extended by 25%; the time constraint in the task configuration parameters is shortened by 15%, and the target area is reduced by 30%. Otherwise, these parameters are relaxed accordingly.
[0034] The adjusted scene structure is instantiated into an interactive training environment. The instantiation process uses a three-dimensional engine to build a physical scene and configures an agent interaction interface with four main interfaces: the state observation interface returns a 128-dimensional vector of the current environmental state; the action execution interface receives 12 discrete actions and 3 continuous action parameters; the communication interaction interface supports 256-bit message transmission between agents; the reward feedback interface calculates a reward signal in the range of -10 to +10 in real time.
[0035] The task evaluation module adjusts the evaluation criteria based on the scene difficulty coefficient. When the difficulty coefficient is 0.3, passing the evaluation can be obtained by completing the basic goals; when the difficulty coefficient reaches 0.7, it is required to complete all goals within the time limit and achieve a 90% collaboration efficiency to obtain a passing evaluation. The evaluation process records the team performance data every 100 milliseconds for subsequent dynamic difficulty adjustment and training effect analysis.
[0036] Figure 2 For the bar chart comparing the task completion rates of the scene difficulty adaptive method in the embodiments of the present invention: This figure shows the performance comparison of three solutions (the present technical solution, the fixed-difficulty solution, and the random-difficulty method) under four different difficulty scenarios. In the simple scenario, all three solutions achieved relatively high performance, reaching 95.3%, 93.5%, and 91.8% respectively, with relatively small differences. In the medium scenario, the performance began to show obvious divergence. The present technical solution maintained a high level of 92.1%, while the fixed-difficulty solution dropped to 84.2%, and the random-difficulty method further decreased to 79.5%. In the complex scenario, the performance difference further widened. The present technical solution still maintained a good performance of 87.4%, while the fixed-difficulty solution dropped to 68.7%, and the random-difficulty method was only 63.2%. In the extreme scenario, the performance difference was the most significant. Although the present technical solution decreased slightly, it still maintained a performance level of 82.9%, while the fixed-difficulty solution and the random-difficulty method dropped significantly to 53.1% and 48.6% respectively. The data clearly shows that as the scenario difficulty increases, the present technical solution demonstrates significant performance advantages and stronger robustness, especially its adaptability in high-difficulty scenarios far exceeds the other two solutions.
[0037] In an optional implementation manner, a multi-objective optimization function is constructed based on the target scenario feature vector, and the feature mapping losses of the environmental feature component, the role feature component, and the task feature component are respectively calculated according to the multi-objective optimization function, including: A feature component representation model is constructed, and based on the feature component representation model, the environmental feature component, the role feature component, and the task feature component are respectively encoded into feature vectors to generate an initial feature vector set; Based on the initial feature vector set, a multi-objective mapping loss function is constructed, and the multi-objective mapping loss function performs weighted calculation on the environmental feature mapping loss, the role feature mapping loss, and the task feature mapping loss to generate a basic mapping loss value; An eigen-causal graph structure is established by using the initial feature vector set. The nodes of the eigen-causal graph structure are the feature vectors in the initial feature vector set, and the edges of the eigen-causal graph structure represent the causal relationships between the feature vectors. The causal strength between the feature nodes is calculated through causal intervention operations; The causal strength is fused with the basic mapping loss value to construct a causally enhanced overall loss function; the parameter gradient is calculated based on the overall loss function, and the parameters of the feature component representation model are updated based on the parameter gradient; Based on the updated feature component representation model, the corresponding weight coefficients of the environmental feature mapping loss, the role feature mapping loss, and the task feature mapping loss are calculated to obtain the feature mapping losses of each feature component.
[0038] Construct a feature component representation model, which includes three encoder modules, namely the environmental feature encoder, the role feature encoder, and the task feature encoder. The environmental feature encoder receives the input environmental features, such as meteorological conditions, lighting conditions, terrain information, etc., and converts these features into a 64-dimensional environmental feature vector through a multi-layer neural network.
[0039] The role feature encoder receives the input role features, such as role type, behavior pattern, interaction preference, etc., and converts these features into a 32-dimensional role feature vector through a convolutional neural network and a fully connected layer. The task feature encoder receives the input task features, such as task difficulty, completion time limit, resource requirements, etc., and converts these features into a 48-dimensional task feature vector through a recurrent neural network. These three feature vectors together form the initial feature vector set.
[0040] For example, for a specific scenario, the environmental features include "indoor space, sufficient light, low noise level", the role features include "adult user, medium technical proficiency, concentrated attention", and the task features include "medium difficulty, need 30 minutes to complete, need medium precision". These features are converted into corresponding feature vectors through their respective encoders.
[0041] Based on the initial feature vector set, construct a multi-objective mapping loss function. This function calculates the weighted sum of the three feature mapping losses to generate the basic mapping loss value. The calculation method of the environmental feature mapping loss is the Euclidean distance between the environmental feature vector and the target scenario feature vector in the environmental dimension; the calculation method of the role feature mapping loss is the negative value of the cosine similarity between the role feature vector and the target scenario feature vector in the role dimension; the calculation method of the task feature mapping loss is the cross-entropy between the task feature vector and the target scenario feature vector in the task dimension. These three losses are combined through weight coefficients, and the initial weights are set to 0.35, 0.35, and 0.3 respectively, to obtain the basic mapping loss value.
[0042] Use the initial feature vector set to establish a feature causal graph structure. In this graph structure, the nodes are the feature vectors in the initial feature vector set, and the edges represent the causal relationships between the feature vectors. The determination of the causal relationships is achieved by analyzing the mutual information and conditional independence between the feature vectors. Specifically, the historical samples of the feature vectors are collected through a sliding window method, the mutual information matrix between the samples is calculated, and the PC algorithm is used to determine the direction and existence of the edges. For a simple scenario, the possible causal graph may include causal links such as "environmental feature → role feature", "environmental feature → task feature", and "role feature → task feature".
[0043] Calculate the causal strength between feature nodes through causal intervention operations. The causal intervention operation is achieved by changing the feature values of the source node and observing the change amplitude of the target node. For example, change "sufficient light" in the environmental features to "insufficient light", and observe the change in the character feature vector. If the change is significant, it indicates that the environmental features have a strong causal impact on the character features. Through multiple intervention experiments, a causal strength matrix is calculated, and each element in the matrix represents the causal strength from one feature to another.
[0044] Fuse the causal strength with the basic mapping loss value to construct an overall loss function with causal enhancement. The specific fusion method is to add the causal strength as a regularization term to the basic mapping loss. If the causal strength between two features is high, a corresponding penalty term is added to the loss function, enabling the model to learn the causal relationship between the features. For example, if the causal strength of the environmental features on the task features is 0.7 (relatively strong), then an environmental-task feature mapping loss term with a weight of 0.7 is added to the overall loss function.
[0045] Calculate the parameter gradients based on the overall loss function and use the gradient descent method to update the parameters of the feature component representation model. In each training batch, collect a set of environmental-character-task feature samples, calculate the overall loss through forward propagation, then calculate the gradients through backpropagation, and update the model parameters. The learning rate is set to 0.001, the Adam optimizer is used for parameter update, the training batch size is 64, and the number of training epochs is 100.
[0046] Based on the updated feature component representation model, calculate the corresponding weight coefficients for the three feature mapping losses. The calculation of the weight coefficients is based on the contribution degree of each feature to the overall loss, which is specifically determined through sensitivity analysis. Make a small perturbation to each feature and observe the change amplitude of the overall loss. The larger the change amplitude, the greater the contribution of the feature to the overall loss, and the larger the corresponding weight coefficient. After multiple iterative calculations, the weight coefficient of the environmental feature mapping loss is finally 0.4, the weight coefficient of the character feature mapping loss is 0.35, and the weight coefficient of the task feature mapping loss is 0.25. These weight coefficients reflect the importance of different feature components in the current scenario.
[0047] Through the above method, a technical solution for constructing a multi-objective optimization function based on the target scenario feature vector and calculating the feature mapping losses for the environmental, character, and task feature components respectively according to this optimization function is realized, providing an effective solution for the target representation of scene adaptation.
[0048] Figure 3 This is a bar chart comparing the multi-objective optimization effects of the feature components in the embodiments of the present invention: The figure shows the performance comparison of three different models (basic feature representation model, multi-object mapping model, causal enhancement mapping model) in five different scenarios. In the complex environment scenario, the performance levels of the three models reach 73.5%, 81.3% and 88.7% respectively, and the causal enhancement mapping model shows the most prominent performance; in the multi-role collaboration scenario, the performances are 68.2%, 77.6% and 85.9% respectively, maintaining a similar performance gap; in the dynamic task scenario, the performances of the three models reach 71.8%, 79.5% and 87.2%, showing the adaptability of the causal enhancement mapping model in the dynamic environment; in the cross-domain fusion scenario, the performances are 65.4%, 74.8% and 82.6% respectively. Although the overall performance has decreased, the relative advantages among the models are still obvious; in the high-pressure scenario, the performances drop to 62.7%, 72.3% and 80.4% respectively, reflecting the impact of the pressure environment on the model performance. Looking at all the test scenarios, the causal enhancement mapping model always maintains a significant performance advantage, being on average about 15 - 20 percentage points higher than the basic feature representation model and about 8 - 10 percentage points higher than the multi-object mapping model, proving the robustness and generalization ability of this model in complex and changeable environments.
[0049] In an optional implementation manner, generating an agent in the scenario reproduction training environment, and training the initial behavior pattern of the agent through a deep reinforcement learning model based on the historical behavior data in the team collaboration knowledge base includes: Generating an agent in the scenario reproduction training environment, extracting historical behavior data from the team collaboration knowledge base, and constructing a deep reinforcement learning model based on the historical behavior data. The deep reinforcement learning model includes a behavior state space, an action space and a reward function; Using the deep reinforcement learning model to iteratively train the agent, updating the policy network parameters of the agent according to the feedback information of the reward function, and training the initial behavior pattern of the agent.
[0050] Construct a scenario reproduction training environment according to the actual application scenario. This environment includes a multi-agent interaction area, environmental state information and a task objective setting module. N agents are generated in the environment, and each agent has the capabilities of perception, decision-making and execution. The perception ability of the agent includes observing the environmental state, and can obtain information such as the current position, the distribution of surrounding objects, and the positions of other agents; the decision-making ability is composed of a deep neural network, which receives the perception information and outputs action instructions; the execution ability then converts the decision into actual actions in the environment. For example, in the warehousing and logistics scenario, 10 handling robots can be generated as agents, and each agent is equipped with a position sensor, a lidar and a communication module, and can perceive obstacles, other agents and target goods in the environment.
[0051] The system accesses a pre-built team collaboration knowledge base, which contains the behavior data of various agents during the execution of historical tasks, such as position trajectories, action selections, task assignment information, etc. In the example of the warehousing and logistics scenario, the knowledge base contains the behavior records of the handling robots collaborating to complete 5000 tasks in the past 3 months, including 500,000 trajectory data points. Each data point records information such as the agent ID, timestamp, position coordinates, current task being executed, power status, and load condition.
[0052] Extract representative historical behavior data from the knowledge base, including regular task execution data and special case handling data. The system uses data cleaning algorithms to remove outliers and redundant data, and retains valid training samples. In the example, the system filters out 4000 task data as valid samples, and eliminates abnormal behavior records and duplicate records caused by equipment failures.
[0053] Based on the filtered historical behavior data, a deep reinforcement learning model is constructed. This model defines three core elements: the behavior state space, the action space, and the reward function. The behavior state space is a set of environmental information that the agent can perceive, including the agent's own state (position, orientation, speed, etc.), the environmental state (obstacle distribution, target position, etc.), and the states of other agents (position, task status, etc.). In the example, the dimension of the state space is 100, including the robot's current two-dimensional coordinates (x, y), orientation angle, linear velocity and angular velocity, battery percentage, distances to obstacles in 8 directions within 10 meters around, the weight of the currently carried goods, the coordinates of the target shelf, and the positions and status information of 5 other nearby robots.
[0054] The action space defines all the behaviors that the agent can execute, and is divided into two types: the discrete action space and the continuous action space. The discrete action space is suitable for scenarios with a limited number of behavior choices, and the continuous action space is suitable for scenarios that require precise control. In the example, the action space of the robot includes 8 discrete moving directions (front, back, left, right, and four diagonal directions) and 3 operation actions (grab, drop, standby), for a total of 11 basic actions.
[0055] The design of the reward function is the key to reinforcement learning, which is used to evaluate the quality of the agent's behavior and guide the learning process. The reward function comprehensively considers factors such as task completion, resource utilization efficiency, collaboration effect, and safety. In the example, the reward function is set as follows: getting a reward of +100 for successfully delivering the goods to the target position; consuming a reward of -0.1 for each time step to encourage quick task completion; getting a reward of -50 for colliding with other robots; getting a reward of -30 for colliding with static obstacles; getting a reward of +0.5 for maintaining an appropriate distance (0.5 - 1.5 meters) from other robots to promote collaboration; getting a reward of +0.2 for each step of driving along the shortest path.
[0056] Based on the above definitions, a deep reinforcement learning network architecture is constructed, including a policy network and a value network. The policy network takes the current state as input and outputs the action probability distribution; the value network takes the current state as input and outputs the state value estimate. The policy network adopts a 4-layer fully connected neural network structure. The number of neurons in the input layer is the same as the dimension of the state space, which is 100. Each of the two hidden layers contains 128 and 64 neurons respectively. The number of neurons in the output layer is the same as the dimension of the action space, which is 11. The ReLU activation function and the Softmax output layer are used. The value network adopts a 3-layer fully connected neural network. The input layer also has 100 neurons. The hidden layer contains 64 neurons. The output layer has 1 neuron representing the state value.
[0057] The constructed deep reinforcement learning model is used to iteratively train the agent. The training process adopts the batch training method. Each batch contains 256 state-action-reward-new state transition samples. In the example, the system sets the total number of training rounds to 5000 rounds, and each round contains 1000 environmental interaction steps. The agent executes actions in the environment, and the environment feedbacks the new state and reward. The agent records these transitions in the experience pool. When the data volume in the experience pool reaches 10,000, the system starts to randomly extract batches for network parameter update.
[0058] During the parameter update process, the system calculates the policy loss and the value loss. The policy loss measures the gap between the current policy and the optimal policy, and the value loss measures the gap between the state value prediction and the actual value. The optimization algorithm uses the Adam optimizer. The learning rate is set to 0.0003, and the momentum parameter is 0.9. The learning rate decays by 10% every 50 rounds. In the example, the loss function values of the policy network and the value network are 4.2 and 8.5 respectively at the beginning of training, and drop to 0.7 and 1.3 after 5000 rounds of training, indicating that the network parameters gradually converge.
[0059] During the training process, the system evaluates the performance of the current policy every 100 rounds. The evaluation metrics include the average task completion time, success rate, number of collisions, and cooperation efficiency. In the example scenario, the average task completion time of the agent under the initial policy is 120 seconds, the success rate is 65%, and the average number of collisions per task is 3.2 times; after 5000 rounds of training, the average task completion time drops to 72 seconds, the success rate increases to 93%, and the average number of collisions decreases to 0.5 times.
[0060] The finally trained agent has the initial behavior pattern and can make cooperative decisions based on the perception information. In the actual warehouse environment, 10 agents can cooperate to complete the multi-target cargo handling task, avoid obstacles autonomously and optimize the path. The task execution efficiency is increased by 62% compared with manual operation and 25% compared with traditional algorithms. The initial behavior pattern of the agent will be used as the basis for subsequent personalized training to further adapt to the specific needs of different scenarios.
[0061] In an alternative embodiment, based on the collaborative behavior data, the impact degree of team collaborative behavior on the task result is analyzed by a counterfactual reasoning method, and a causal relationship graph between collaborative behavior and task performance is constructed. The causal relationship graph dynamically quantifies the intensity of the synergy effect among team members. Combining the temporal attention mechanism to track the ability evolution trajectory includes: Construct a collaborative behavior vector and a task result vector based on the collaborative behavior data, and use a causal intervention operator to calculate the causal effect of each behavior feature in the collaborative behavior vector on each result index in the task result vector. The causal effect characterizes the impact intensity of the collaborative behavior on the task result, and construct a causal relationship graph between team collaborative behavior and task performance based on the causal effect; Calculate the synergy intensity among team members based on the causal relationship graph, combine the temporal weight and the attention weight to weight the causal effect, and obtain the dynamic synergy effect intensity; Construct a team ability phase space based on the dynamic synergy effect intensity, calculate an ability index based on the team ability phase space. The ability index is used to characterize the stability of the team ability evolution, and input the ability index into a multi-scale temporal attention mechanism; The multi-scale temporal attention mechanism calculates the dynamic correlation weight of temporal features using a query matrix, a key matrix, and a value matrix, and constructs a Jacobian matrix based on the dynamic correlation weight. The Jacobian matrix describes the local linear features of the ability phase space; Combine the ability index and the Jacobian matrix to calculate a bifurcation discriminant, predict the transition probability of the team ability based on the bifurcation discriminant. The transition probability indicates the turning point of the ability evolution; Combine the dynamic correlation weight and the transition probability to track and predict the team ability evolution trajectory.
[0062] As Figure 4 shown, the method includes: Collect the collaborative behavior data generated by team members during the task execution. These data can include indicators such as the communication frequency, response time, task assignment, and decision-making participation among members. Data collection can be obtained through collaborative platform log records, sensor monitoring, or behavior annotation methods in specific tasks. For example, during the project development of a 5-person team, record data points such as the number of code submissions, the number of review comments, and the contribution degree of problem-solving for each member.
[0063] Convert the collected collaborative behavior data into a structured collaborative behavior vector, and at the same time construct evaluation indicators such as task completion quality, efficiency, and innovation into a task result vector. For example, the collaborative behavior vector can include [communication frequency: 0.8, resource sharing degree: 0.6, decision-making participation degree: 0.7], and the task result vector can include [completion quality: 0.85, time efficiency: 0.75, innovation index: 0.6] and other dimensions. The system calculates the influence degree of each behavior feature on the result index through the causal intervention operator. Specifically, during implementation, counterfactual intervention is performed on a certain collaborative behavior feature, such as changing the "communication frequency" from 0.8 to 0.4, observing the change amplitude of each index in the task result vector, and repeating this process to establish a causal effect matrix. Calculations show that the causal effect of communication frequency on completion quality is 0.35, and the causal effect on time efficiency is 0.28. The higher the value, the more significant the influence.
[0064] Based on the causal effect matrix, construct a causal relationship diagram of team collaborative behavior and task performance. The nodes in the diagram represent collaborative behavior features or task result indicators, and the weights of the edges represent the size of the causal effect. In practical applications, visually display this causal relationship diagram, showing that the influence weight of "resource sharing degree" on "completion quality" is 0.42, and the influence weight on "innovation index" is 0.56, helping managers intuitively understand the correlation strength between behavior and performance.
[0065] Calculate the synergy strength among team members. First, identify the collaborative interaction patterns among members, such as the collaboration frequency, complementarity, and synergy effect between member A and member B. Combine the importance weights of time-series data with the attention weights of specific behaviors to dynamically weight the aforementioned causal effects. For example, in the early stage of a project, the weight of communication and collaboration may be 0.65, while in the later stage of the project, the weight of execution ability may increase to 0.72. The calculation results show that over time, the synergy effect between team members A and B in problem-solving has increased from 0.56 to 0.78, indicating that their collaboration pattern is becoming increasingly mature.
[0066] Construct a team ability phase space based on the dynamic synergy strength. This space describes the state changes of the team on multi-dimensional ability indicators. In the phase space, calculate the ability index to characterize the stability of the team ability evolution. Specifically, select key ability dimensions such as problem-solving ability, innovation ability, and adaptability to construct the phase space, and track the dynamic change trajectories of the team state on these dimensions. Calculations show that the ability index of a certain team is 0.12, indicating that its ability evolution is relatively stable; while the index of another team reaches 0.45, indicating that its ability state fluctuates greatly and may require management intervention.
[0067] Input the calculated capability index into the multi-scale time-series attention mechanism. This mechanism calculates the dynamic correlation weights of time-series features through a query matrix, a key matrix, and a value matrix. In implementation, the query matrix represents the capability state that needs attention currently, the key matrix stores the historical capability state features, and the value matrix contains the corresponding capability performance data. Through the interactive operations of these three matrices, the correlation weights of capability features at different time points are obtained. For example, the influence weight of the recent collaboration mode on the current innovation capability is 0.63, and the influence weight of the early training effect on the current problem-solving capability is 0.24.
[0068] Construct a Jacobian matrix based on the dynamic correlation weights. This matrix describes the mutual influence relationship between each dimension in the capability phase space. The matrix element value represents the immediate influence intensity of a change in one capability dimension on another dimension. Practical applications show that the influence coefficient of the team collaboration capability on the problem-solving efficiency is 0.58, while the influence coefficient of the information sharing level on the innovation capability is 0.71.
[0069] Combine the capability index and the Jacobian matrix to calculate a bifurcation discriminant, which is used to detect the conditions under which the capability state may undergo a qualitative change. When the value of the discriminant exceeds a preset threshold (such as 0.85), it indicates that the team's capability may experience a leap. Calculate the leap probability of the team's capability based on the discriminant. For example, after a specific team training, the system calculates that the leap probability of the team collaboration capability improving in the next stage is 0.72, indicating that the training is likely to produce a positive effect.
[0070] Combine the dynamic correlation weights and the leap probability to construct a prediction model for the evolution of the team's capability. This model can predict the development trajectory of the team's capability in a future period and identify possible turning points. For example, the prediction shows that the innovation capability of the team will increase steadily by about 15% in the next quarter, but there may be a 30% opportunity for a capability leap when the cross-departmental project starts. These prediction results provide accurate intervention timing and direction guidance for team managers.
[0071] In an alternative implementation manner, combining the capability index and the Jacobian matrix to calculate a bifurcation discriminant, and predicting the leap probability of the team's capability based on the bifurcation discriminant includes: Construct a bifurcation parameter vector by combining the capability index and the eigenvalues of the Jacobian matrix. The bifurcation parameter vector contains the maximum real part of the capability index and the eigenvalues. Based on the bifurcation parameter vector, calculate the bifurcation discriminant through the determinant of the Jacobian matrix and the exponential smoothing term. The bifurcation discriminant is used to detect the mutation characteristics of the team's capability state. Calculate the critical value of the bifurcation discriminant within the observation time window. The critical value represents the threshold condition for the team's capability state to undergo a leap. Input the difference between the bifurcation discriminant and the critical value into a non-linear mapping function. The non-linear mapping function is constructed based on the sigmoid function to calculate the transition probability of the team's ability, and the transition probability characterizes the possibility of a mutation in the team's ability. Construct a confidence interval for the transition probability. The confidence interval calculates the upper and lower bounds through an uncertainty function, and the uncertainty function is dynamically adjusted according to the change of the bifurcation discriminant to achieve accurate prediction of the team's ability transition.
[0072] Obtain the time-series data during the collaboration of team members, and represent the various ability indicators of the team through the system state vector. For a R & D team composed of 5 members, 30 days of continuous collaboration data are collected, including indicators such as communication frequency, task completion rate, knowledge contribution, problem-solving time, and the number of innovation proposals. The system calculates the ability index and the Jacobian matrix for each time point through the adaptive sliding window technology.
[0073] The calculation of the ability index adopts the phase space reconstruction method, selects parameters with an embedding dimension of 4 and a time delay of 2, and analyzes the dynamic time series of team collaboration through the nearest neighbor orbit tracking algorithm. For example, for the task completion rate indicator, its ability index is calculated to be 0.127, indicating a certain degree of instability in this indicator.
[0074] The calculation of the Jacobian matrix is based on the dynamic model of the team's ability, and this model includes the coupling relationship of five key variables. Through the numerical differentiation method, the influence of state variables on the system evolution is evaluated at each time point. For example, in the observation on the 15th day, the calculated Jacobian matrix contains a quantitative description of the interaction effect among team members, where the influence coefficient of knowledge contribution on problem-solving time is -0.42, meaning that knowledge sharing can significantly reduce problem-solving time.
[0075] When constructing the bifurcation parameter vector, the system combines the ability index with the eigenvalues of the Jacobian matrix. Specifically, five eigenvalues of the Jacobian matrix are extracted, and the eigenvalue with the largest real part is selected and combined with the ability index to form a two-dimensional parameter vector. For example, in the data on the 20th day, the largest real part of the eigenvalue is 0.086, which is combined with the ability index 0.127 to form the bifurcation parameter vector [0.127, 0.086].
[0076] The calculation of the bifurcation discriminant combines the determinant information of the Jacobian matrix and the exponential smoothing term. The system first calculates the determinant value of the Jacobian matrix to get -0.0032, then introduces a smoothing factor of 0.85 to construct an exponential smoothing term of 0.0241. The bifurcation discriminant is obtained through the weighted combination of these two terms as 0.0209, and this value is used to detect whether the team's ability is close to the bifurcation point. The calculation of the discriminant is dynamically updated within a 10-day observation window to form a time series of the discriminant.
[0077] The determination of the critical value is completed through the statistical analysis of historical data. The system analyzed the records of the team's ability changes in the past six months, identified 20 ability transition events, and extracted the discriminant values before each transition. Through percentile analysis, the critical value was determined to be 0.025, meaning that when the discriminant exceeds this value, there is a high probability that the team's ability will transition.
[0078] The transition probability is calculated through a non-linear mapping function, which is constructed based on the sigmoid function, with the steepness parameter set to 15 and the offset parameter set to 0.5. Specifically, when the bifurcation discriminant is 0.0209 and the critical value is 0.025, the difference between the two is -0.0041. After inputting into the non-linear mapping function, the transition probability obtained is 0.437, indicating that there is a 43.7% probability that the team's ability will change significantly in the near future.
[0079] A confidence interval for the transition probability is also constructed and dynamically adjusted through an uncertainty function. This uncertainty function takes into account data noise, model parameters, and external interference factors, providing a narrower confidence interval when the discriminant is close to the critical value and a wider interval when it is far from the critical value. For example, when the transition probability is 0.437, the upper and lower bounds of the 95% confidence interval are 0.512 and 0.362 respectively.
[0080] Continuously monitor the changes in the indicators. When the transition probability exceeds 0.6 for three consecutive days, trigger the early warning mechanism. For example, in the middle of the project, the system detected a decrease in communication frequency and fluctuations in task completion rate, resulting in a transition probability of 0.78, successfully predicting the subsequent decline in team collaboration efficiency.
[0081] The advantage of this method is that it can perceive potential mutations in the team's ability in advance, providing a time window for managers to intervene. Experimental verification shows that in 30 test cases, the prediction accuracy of this method reaches 84%, and the average early warning time is 3.7 days, improving the accuracy by 26% and the early warning time by 1.5 days compared with traditional statistical methods.
[0082] Figure 5 The following is a bar chart for the analysis and comparison of the transition probabilities of different team abilities in the embodiments of the present invention: This figure shows the performance comparison of three different methods (basic classification method, Lyapunov enhancement method, and dynamic confidence interval method) under four test scenarios. In the standard scenario test, the performances of the three methods are 67.8%, 78.4%, and 83.6% respectively, and the dynamic confidence interval method performs the best; in the high-pressure environment test, the performances are 59.3%, 72.6%, and 79.2% respectively, and the performances of all methods decline but maintain a relative advantage relationship; in the team collaboration test, the three methods reach relatively high performance levels of 74.2%, 85.1%, and 88.7%, showing good adaptability in the collaboration scenario; in the long-term project test, the performances are 62.5%, 76.9%, and 84.3% respectively, indicating that the dynamic confidence interval method has a more stable performance in continuous tasks. Overall, the dynamic confidence interval method maintains a significant performance advantage in all test scenarios, leading the basic classification method by about 20 percentage points on average and leading the Lyapunov enhancement method by about 5 - 8 percentage points, demonstrating the robustness and adaptability of this method in different application scenarios.
[0083] In an alternative implementation, according to the ability evolution trajectory, a team collaboration assessment report is generated; based on the team collaboration assessment report, a team collaboration optimization plan is output, and the parameter configuration of the scenario reproduction training environment is adjusted specifically, and updating the training strategy of the multi-agent collaborative learning network includes: The ability evolution trajectory includes multi-dimensional ability indicators. Based on the multi-dimensional ability indicators, the individual ability growth rate and the team collaboration efficiency are calculated. The individual ability growth rate represents the speed of ability improvement, and the team collaboration efficiency represents the collaboration effect; the individual ability growth rate and the team collaboration efficiency are weighted and combined to obtain a team collaboration comprehensive score, and the team collaboration comprehensive score reflects the overall performance of the team; Based on the team collaboration comprehensive score, a multi-dimensional evaluation system is constructed. The multi-dimensional evaluation system includes an individual dimension, a team dimension, and a task dimension, and a team collaboration assessment report is generated; according to the team collaboration assessment report, an optimization objective function is constructed. The optimization objective function combines performance loss and optimization cost, and a team collaboration optimization plan is generated based on resource constraints, time constraints, and effect constraints; The team collaboration optimization plan is converted into the parameter configuration of the scenario reproduction training environment. The parameter configuration includes training difficulty parameters and scenario complexity parameters. The scenario reproduction training environment is dynamically adjusted, and the training strategy of the multi-agent collaborative learning network is updated based on the parameter configuration.
[0084] By analyzing the ability evolution trajectory, a team collaboration assessment report is generated, and based on this, a team collaboration optimization plan is output, the parameter configuration of the scenario reproduction training environment is adjusted specifically, and the training strategy of the multi-agent collaborative learning network is updated.
[0085] The ability evolution trajectory includes multi-dimensional ability indicators, which cover five dimensions: communication ability, decision-making efficiency, resource allocation ability, problem-solving ability, and adaptability. Each dimension uses a standardized score from 0 to 100, which is formed by collecting the performance data of the agent during the training process. For example, the ability indicators of three agents A, B, and C in a team at the initial stage of training are: A(65,70,60,75,80), B(80,65,75,60,70), C(70,75,65,80,60), and after 50 rounds of training, they evolve into: A(80,85,75,85,90), B(90,80,85,75,85), C(85,90,80,90,75).
[0086] Based on the multi-dimensional ability indicators, calculate the individual ability growth rate and the team collaboration efficiency. The individual ability growth rate is calculated by dividing the difference between the current ability value and the benchmark ability value by the number of training rounds. Taking agent A as an example, the growth rate of its communication ability is (80 - 65) / 50 = 0.3. The team collaboration efficiency is calculated by the weighted average of three indicators: team task completion rate, collaborative conflict resolution rate, and information sharing efficiency. In practical applications, the weight of the task completion rate can be set to 0.5, the weight of the collaborative conflict resolution rate can be set to 0.3, and the weight of the information sharing efficiency can be set to 0.2. Assuming that the three indicators of a team in a specific training stage are 85%, 90%, and 80% respectively, then the team collaboration efficiency is 85%×0.5 + 90%×0.3 + 80%×0.2 = 85.5%.
[0087] Weightedly combine the individual ability growth rate and the team collaboration efficiency to obtain the comprehensive team collaboration score. The weighting coefficients can be adjusted according to the specific application scenario. A typical configuration is that the weight of the individual ability growth rate is 0.4 and the weight of the team collaboration efficiency is 0.6. If the average individual ability growth rate of a team is 0.28 and the team collaboration efficiency is 85.5%, then the comprehensive team collaboration score is 0.28×0.4 + 0.855×0.6 = 0.625. This score reflects the overall performance of the team, ranging from 0 to 1, and the closer it is to 1, the better the team collaboration effect.
[0088] Construct a multi-dimensional evaluation system based on the comprehensive score of team collaboration, including individual dimension, team dimension, and task dimension. The individual dimension focuses on the ability development of each agent, recording the growth curve and bottleneck points of each dimension; the team dimension analyzes the team structure, role allocation, and interaction patterns; the task dimension evaluates the performance differences of the team in tasks of different difficulties. The evaluation system associates the indicators of each dimension with the comprehensive score of team collaboration through a weight matrix to generate a team collaboration evaluation report. This report includes four parts: a radar chart of team ability distribution, collaborative mode analysis, strength and weakness analysis, and development trend prediction. For example, the evaluation results of a certain team show high communication efficiency (90 points) but weak resource allocation ability (65 points), and the collaborative mode analysis finds that the collaborative quality of this team decreases by 30% in a high-pressure environment.
[0089] Construct an optimization objective function based on the team collaboration evaluation report. This function combines performance loss and optimization cost to generate a team collaboration optimization plan based on resource constraints, time constraints, and effect constraints. Performance loss is quantified by the gap between the current performance and the ideal performance, and the optimization cost includes computing resource consumption, training time, and model complexity. The constraint conditions are set as follows: the computing resources do not exceed 120% of the budget, the training time does not exceed 72 hours, and the performance improvement is not less than 15%. The optimization plan includes three aspects: ability improvement strategy, collaborative mechanism optimization, and task adaptation adjustment, and proposes improvement measures for the deficiencies found in the evaluation report. For example, for the problem of weak resource allocation ability, the proportion of resource competition training scenarios can be increased from 20% to 35%, and a resource sharing reward mechanism can be introduced, with the reward coefficient set to 1.5 times the basic reward.
[0090] Convert the team collaboration optimization plan into parameter configurations for the scenario reproduction training environment, including training difficulty parameters and scenario complexity parameters. The training difficulty parameters control the task challenge, covering three aspects: time pressure, resource limitation, and interference factors; the scenario complexity parameters define the environmental change frequency, information uncertainty, and decision-making constraint conditions. For example, for the need to improve decision-making efficiency, the time pressure parameter can be adjusted from 0.6 to 0.8, and at the same time, the information uncertainty parameter can be reduced from 0.5 to 0.3 to train the ability of agents to make more accurate decisions within limited time. In addition, based on the parameter configuration, update the training strategies of the multi-agent collaborative learning network, including adjusting the learning rate decay curve, reward function weight, and experience replay sampling strategy. For communication ability training, the relevant reward weight can be increased from 0.2 to 0.35, and a communication failure penalty mechanism can be added, with the penalty coefficient set to 0.15.
[0091] By dynamically adjusting the scenario reproduction training environment and updating the training strategy, the collaboration ability of the agent team is continuously optimized. In a collaborative navigation task involving 5 agents, after applying this method, the comprehensive team collaboration score has increased from 0.625 to 0.782, the task completion time has been shortened by 21.5%, and the resource utilization efficiency has been improved by 18.3%, significantly verifying the effectiveness of this method.
[0092] In the second aspect of the embodiments of the present invention, a team collaboration training system based on scenario reproduction and multi-agent collaboration is provided, including: A first unit, configured to establish a team collaboration knowledge base based on pre-acquired team collaboration historical data, construct a multi-agent interaction scenario according to the team collaboration knowledge base, and map virtual environment parameters, agent role information, and task objective information in the team collaboration knowledge base to the multi-agent interaction scenario to generate a scenario reproduction training environment; A second unit, configured to generate agents in the scenario reproduction training environment, and train the initial behavior patterns of the agents through a deep reinforcement learning model based on the historical behavior data in the team collaboration knowledge base; A third unit, configured to construct a multi-agent collaborative learning network enhanced by knowledge distillation, and feedback the collaborative behavior data generated by team members interacting with the agents in the scenario reproduction training environment to the multi-agent collaborative learning network in real time; A fourth unit, configured to analyze the influence degree of team collaboration behavior on task results through a counterfactual reasoning method based on the collaborative behavior data, construct a causal relationship graph between collaborative behavior and task performance, dynamically quantify the synergy strength between team members in the causal relationship graph, and track the ability evolution trajectory in combination with the temporal attention mechanism; A fifth unit, configured to generate a team collaboration evaluation report according to the ability evolution trajectory; based on the team collaboration evaluation report, output a team collaboration optimization plan, targetedly adjust the parameter configuration of the scenario reproduction training environment, and update the training strategy of the multi-agent collaborative learning network.
[0093] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0094] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0095] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A team collaboration training method based on scenario reproduction and multi-agent collaboration, characterized in that, include: A team collaboration knowledge base is established based on pre-acquired team collaboration historical data, a multi-agent interaction scenario is constructed according to the team collaboration knowledge base, virtual environment parameters, agent role information, and task target information in the team collaboration knowledge base are mapped to the multi-agent interaction scenario, and a scenario reproduction training environment is generated; Generate an intelligent agent in the scenario reproduction training environment, and train the initial behavior pattern of the intelligent agent through a deep reinforcement learning model based on the historical behavior data in the team collaboration knowledge base; Construct a multi-agent collaborative learning network enhanced by knowledge distillation, and feed back the collaborative behavior data generated by the interaction between the team members and the agent in the scene reproduction training environment to the multi-agent collaborative learning network in real time; Based on the collaborative behavior data, the influence of team collaborative behavior on task results is analyzed by counterfactual reasoning method, and a causal relationship diagram between collaborative behavior and task performance is constructed. The causal relationship diagram dynamically quantifies the intensity of synergy between team members, and combines the temporal attention mechanism to track the evolution trajectory of capabilities; According to the capability evolution trajectory, a team collaboration evaluation report is generated; based on the team collaboration evaluation report, a team collaboration optimization plan is output, the parameter configuration of the scenario reproduction training environment is adjusted in a targeted manner, and the training strategy of the multi-agent collaborative learning network is updated.
2. The method according to claim 1, wherein Mapping the virtual environment parameters, agent role information, and task target information in the team collaboration history data to the multi-agent interaction scene to generate a scene reproduction training environment includes: Constructing a feature vector space mapping model, extracting features of virtual environment parameters, agent role information, and task target information based on the feature vector space mapping model, and generating a target scene feature vector, wherein the target scene feature vector includes an environment feature component, a role feature component, and a task feature component; Constructing a multi-objective optimization function based on the target scene feature vector, calculating feature mapping losses for the environment feature component, the role feature component, and the task feature component according to the multi-objective optimization function, and optimizing the target scene feature vector based on a weighted result of the feature mapping losses; The optimized target scene feature vector is input into the generator of the generative adversarial network, and the generator reconstructs the target scene feature vector through a multi-layer transposed convolutional network to generate an initial scene structure including environment layout parameters, role attribute parameters and task configuration parameters; Building a scene difficulty adaptive adjustment module based on the initial scene structure, the scene difficulty adaptive adjustment module calculates a scene difficulty coefficient according to a preset performance indicator, and dynamically adjusts the initial scene structure based on the scene difficulty coefficient; The adjusted initial scene structure is instantiated as a training environment, and an agent interaction interface and a task evaluation module are configured in the training environment. The task evaluation module dynamically adjusts the evaluation criteria based on the scene difficulty coefficient to achieve adaptive generation of a scene reproduction training environment.
3. The method according to claim 2, characterized in that Construct a multi-objective optimization function based on the target scenario feature vector, and calculate the feature mapping losses for the environmental feature component, the role feature component, and the task feature component according to the multi-objective optimization function, including: Construct a feature component representation model, and encode the environmental feature component, the role feature component, and the task feature component into feature vectors respectively based on the feature component representation model to generate an initial feature vector set; Construct a multi-objective mapping loss function based on the initial feature vector set. The multi-objective mapping loss function performs weighted calculation on the environmental feature mapping loss, the role feature mapping loss, and the task feature mapping loss to generate a basic mapping loss value; Establish a feature causal graph structure using the initial feature vector set. The nodes of the feature causal graph structure are the feature vectors in the initial feature vector set, and the edges of the feature causal graph structure represent the causal relationships between the feature vectors. Calculate the causal strength between the feature nodes through causal intervention operations; Fuse the causal strength with the basic mapping loss value to construct a causally enhanced overall loss function; calculate the parameter gradients based on the overall loss function, and update the parameters of the feature component representation model based on the parameter gradients; Calculate the corresponding weight coefficients for the environmental feature mapping loss, the role feature mapping loss, and the task feature mapping loss based on the updated feature component representation model to obtain the feature mapping losses of each feature component.
4. The method according to claim 1, characterized in that, Generate an agent in the scenario reproduction training environment, and train the initial behavior pattern of the agent through a deep reinforcement learning model based on the historical behavior data in the team collaboration knowledge base, including: Generate an agent in the scenario reproduction training environment, extract historical behavior data from the team collaboration knowledge base, and construct a deep reinforcement learning model based on the historical behavior data. The deep reinforcement learning model includes a behavior state space, an action space, and a reward function; Use the deep reinforcement learning model to iteratively train the agent, and update the policy network parameters of the agent according to the feedback information of the reward function to train the initial behavior pattern of the agent.
5. The method according to claim 1, characterized in that, Analyze the influence degree of team collaboration behavior on the task result through a counterfactual reasoning method based on the collaborative behavior data, construct a causal relationship graph between collaborative behavior and task performance, and dynamically quantify the synergy strength between team members by combining with the temporal attention mechanism to track the ability evolution trajectory, including: Construct a collaborative behavior vector and a task result vector based on the collaborative behavior data, and calculate the causal effect of each behavior feature in the collaborative behavior vector on each result index in the task result vector using a causal intervention operator. The causal effect represents the influence strength of the collaborative behavior on the task result, and construct a causal relationship graph between team collaboration behavior and task performance based on the causal effect; Calculate the synergy strength between team members based on the causal relationship graph, and combine the temporal weight and the attention weight to weight the causal effect to obtain the dynamic synergy strength; Construct a team ability phase space based on the dynamic synergy strength, calculate an ability index based on the team ability phase space, where the ability index is used to characterize the stability of the evolution of the team ability, and input the ability index into a multi-scale temporal attention mechanism; The multi-scale temporal attention mechanism calculates the dynamic correlation weights of temporal features using a query matrix, a key matrix, and a value matrix, and constructs a Jacobian matrix based on the dynamic correlation weights, where the Jacobian matrix describes the local linear features of the ability phase space; Combine the ability index and the Jacobian matrix to calculate a bifurcation discriminant, predict the transition probability of the team ability based on the bifurcation discriminant, where the transition probability indicates the turning point of the ability evolution; combine the dynamic correlation weights and the transition probability to track and predict the evolution trajectory of the team ability.
6. The method according to claim 5, characterized in that, Combining the ability index and the Jacobian matrix to calculate a bifurcation discriminant, and predicting the transition probability of the team ability based on the bifurcation discriminant includes: Construct a bifurcation parameter vector by combining the ability index and the eigenvalues of the Jacobian matrix, where the bifurcation parameter vector includes the maximum real part of the ability index and the eigenvalues; Calculate a bifurcation discriminant based on the bifurcation parameter vector through the determinant of the Jacobian matrix and an exponential smoothing term, where the bifurcation discriminant is used to detect the mutation characteristics of the team ability state; Calculate the critical value of the bifurcation discriminant within the observation time window, where the critical value represents the threshold condition for the transition of the team ability state, and input the difference between the bifurcation discriminant and the critical value into a non-linear mapping function; The non-linear mapping function is constructed based on the sigmoid function to calculate the transition probability of the team ability, where the transition probability represents the possibility of the mutation of the team ability; Construct a confidence interval for the transition probability, where the upper and lower bounds of the confidence interval are calculated by an uncertainty function, and the uncertainty function is dynamically adjusted with the change of the bifurcation discriminant to achieve accurate prediction of the team ability transition.
7. The method according to claim 1, wherein Generate a team collaboration evaluation report according to the ability evolution trajectory; based on the team collaboration evaluation report, output a team collaboration optimization plan, and specifically adjust the parameter configuration of the scenario reproduction training environment, and update the training strategy of the multi-agent collaborative learning network, including: The ability evolution trajectory includes multi-dimensional ability indicators. Calculate the individual ability growth rate and the team collaboration efficiency based on the multi-dimensional ability indicators. The individual ability growth rate characterizes the ability improvement speed, and the team collaboration efficiency characterizes the collaboration effect; perform a weighted combination of the individual ability growth rate and the team collaboration efficiency to obtain a team collaboration comprehensive score, where the team collaboration comprehensive score reflects the overall performance of the team; Construct a multi-dimensional evaluation system based on the team collaboration comprehensive score, where the multi-dimensional evaluation system includes an individual dimension, a team dimension, and a task dimension, and generate a team collaboration evaluation report; construct an optimization objective function according to the team collaboration evaluation report, where the optimization objective function combines the performance loss and the optimization cost, and generate a team collaboration optimization plan based on resource constraints, time constraints, and effect constraints; The team collaboration optimization scheme is converted into a parameter configuration of the scenario reproduction training environment, wherein the parameter configuration includes a training difficulty parameter and a scenario complexity parameter, the scenario reproduction training environment is dynamically adjusted, and the training strategy of the multi-agent collaborative learning network is updated based on the parameter configuration.
8. A team collaboration training system based on scenario reproduction and multi-agent collaboration, for implementing the method described in any one of the foregoing claims 1-7, characterized in that, include: The first unit is used to establish a team collaboration knowledge base based on pre-acquired team collaboration historical data, construct a multi-agent interaction scene according to the team collaboration knowledge base, map the virtual environment parameters, agent role information and task target information in the team collaboration knowledge base to the multi-agent interaction scene, and generate a scene reproduction training environment; The second unit is used to generate an intelligent agent in the scenario reproduction training environment, and train the initial behavior pattern of the intelligent agent through a deep reinforcement learning model based on the historical behavior data in the team collaboration knowledge base; The third unit is used to construct a multi-agent collaborative learning network enhanced by knowledge distillation, and to feed back the collaborative behavior data generated by the interaction between the team members and the agent in the scene reproduction training environment to the multi-agent collaborative learning network in real time; The fourth unit is used to analyze the influence of team collaboration behavior on task results through counterfactual reasoning method based on the collaboration behavior data, and to construct a causal relationship diagram between collaboration behavior and task performance. The causal relationship diagram dynamically quantifies the intensity of synergy between team members and tracks the evolution trajectory of capabilities in combination with the temporal attention mechanism; The fifth unit is used to generate a team collaboration evaluation report based on the capability evolution trajectory; based on the team collaboration evaluation report, output a team collaboration optimization plan, specifically adjust the parameter configuration of the scenario reproduction training environment, and update the training strategy of the multi-agent collaborative learning network.
9. An electronic device, characterized in that, include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Deep reinforcement learning multi-agent cooperation method based on knowledge distillation
CN113449867A
Multi-agent online learning method based on decision attention mechanism
CN117391153A
Visual interpretation method and system for man-machine collaborative decision
CN118467801A
Robot agent reinforcement learning training method and system in complex scene
CN119129642A
Microscopic three-dimensional simulation method and system for macro-micro integration of composite traffic network
CN119475738A
Cited By
High-degree-of-freedom dexterous hand intelligent control method and system based on collaborative dimension reduction
CN121083620A
High-dof dexterous hand intelligent control method and system based on collaborative dimension reduction
CN121083620B
Evaluation system for college student occupational scene simulation training
CN121437232A