Multi-agent task cooperation method, device and equipment and storage medium

Through the multimodal large language model, the multimodal data of the multimodal system is deeply fusion and semantic encoding, which solves the problems of data fusion and concept drift in multi-agent task collaboration, and improves the system's coordination ability and long-term adaptability.

CN120046100AInactive Publication Date: 2025-05-27SHENZHEN FUTURE QINGYAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 26 Cited by

Patent Information

Application Number
CN202510103530.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing multi-agent task collaboration methods are difficult to effectively process and integrate multimodal data, and cannot cope with concept drift problems in long-term operation.

Method used

The multimodal large language model is used to deeply fusion and unified semantic encoding of multimodal data sets to generate high-dimensional cross-modal feature embedding and semantic representation files for task requirements semantic mapping and task division. At the same time, the concept offset is detected in real time and the model is dynamically adjusted to cope with concept drift.

Benefits of technology

The task coordination capabilities and long-term performance of multi-agent systems are improved, and adaptability in complex dynamic environments is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046100A_ABST
    Figure CN120046100A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent task collaboration method, device and equipment and a storage medium, and the method comprises the steps: carrying out the deep fusion and unified semantic coding of a collected multi-modal data set through a multi-modal large language model, and obtaining a high-dimensional cross-modal feature embedding and semantic representation file, the task requirement mapping module is used for enabling local task requirements of multiple agents to correspond to cross-modal semantics to obtain task requirement semantic mapping, and task division and time arrangement are carried out; when multiple agents execute tasks, key data and operation results are sampled in real time and compared with high-dimensional semantic representation, a concept offset detection result is obtained, and when it is detected that the concept drifts progressively, the multi-modal large language model is dynamically adjusted. According to the method, the communication and cooperation efficiency among multiple agents is enhanced by using a large language model, and task allocation and collaborative decision are optimized through task demand semantic mapping; and concept drift detection and a dynamic adjustment mechanism are introduced, so that the long-term adaptability in a complex dynamic environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of device control, and in particular, to a multi-agent task collaboration method, device, equipment, and storage medium. Background Art

[0002] In the field of multi-agent systems, task collaboration is a key challenge. Traditional multi-agent task collaboration methods are mainly based on predefined rules and simple communication protocols, and it is difficult to cope with complex and dynamic environments and task requirements. With the development of artificial intelligence technology, researchers have begun to explore the use of machine learning methods to improve the collaboration ability of multi-agent systems.

[0003] In recent years, the progress of multi-modal learning technology has brought new opportunities for multi-agent task collaboration. Multi-modal learning can integrate various data types such as vision, text, and speech, so as to obtain a more comprehensive and rich representation of the environment and tasks. However, how to effectively apply multi-modal learning technology to multi-agent systems to achieve efficient task allocation and collaborative decision-making is still an unsolved problem.

[0004] In addition, large language models have made significant breakthroughs in the field of natural language processing, demonstrating powerful understanding and generation capabilities. Some studies have begun to try to introduce large language models into multi-agent systems to improve the communication and collaboration efficiency between agents. However, how to combine large language models with multi-modal learning and achieve multi-agent task collaboration based on this still lacks a systematic solution.

[0005] Another important challenge is the problem of concept drift. In a multi-agent system running for a long time, the dynamic changes of the environment and tasks may cause the performance of the model to gradually decline. Most of the existing methods adopt a fixed model structure and it is difficult to adapt to this progressive concept drift. Therefore, how to detect concept drift in real time and dynamically adjust the model to maintain the long-term performance of the system has become an urgent problem to be solved. Summary of the Invention

[0006] The main purpose of the present invention is to solve the technical problems of the lack of a systematic solution for effectively processing and integrating multi-modal data in existing multi-agent task collaboration methods and the inability to cope with the problem of concept drift during long-term operation; The first aspect of the present invention provides a multi-agent task collaboration method, and the multi-agent task collaboration method includes: Collect visual images, text descriptions, voice instructions, and structured sensor signals, and perform spatio-temporal alignment and format conversion processing on the collected data to obtain a multi-modal data set; Utilize a multi-modal large language model to deeply fuse various modal inputs of the collected multi-modal dataset and perform unified semantic encoding to obtain high-dimensional cross-modal feature embeddings and semantic representation files; Based on the high-dimensional cross-modal feature embeddings and semantic representation files, correspond the local task requirements of multiple agents with cross-modal semantics to obtain a task requirement semantic mapping, and according to the task requirement semantic mapping, perform task division and time arrangement of the multiple agents to obtain a multi-agent cooperation strategy file; When the multiple agents execute the multi-agent cooperation strategy file, perform real-time sampling on the key data and operation results during the execution process, and compare the sampled data with the high-dimensional semantic representation to obtain a concept deviation detection result. When the concept deviation detection result indicates a progressive drift of the concept, perform dynamic adjustment on the multi-modal large language model.

[0007] Optionally, in the first implementation manner of the first aspect of the present invention, the collecting of visual images, text descriptions, voice instructions, and structured sensor signals, and the spatio-temporal alignment and format conversion processing of the collected data to obtain a multi-modal dataset includes: Perform convolutional neural network processing on the visual image to extract an image feature vector, apply a BERT model to the text description for word embedding to obtain a text feature vector, extract audio features from the voice instruction through MFCC transformation to obtain a voice feature vector, and perform wavelet transformation on the structured sensor signal to obtain a sensor feature vector; Align the image feature vector, text feature vector, voice feature vector, and sensor feature vector according to the timestamp to construct a multi-modal time series matrix, and apply the dynamic time warping algorithm to the multi-modal time series matrix to obtain spatio-temporally aligned data; Map the data of different modalities in the spatio-temporally aligned data to a unified numerical range, and reduce the feature dimension through principal component analysis to obtain a reduced-dimensional feature matrix; Apply a tensor decomposition algorithm to the reduced-dimensional feature matrix to extract shared and independent modal features, and integrate them into a tensor structure in a unified format to obtain the multi-modal dataset.

[0008] Optionally, in the second implementation manner of the first aspect of the present invention, the utilizing of a multi-modal large language model to deeply fuse various modal inputs of the collected multi-modal dataset and perform unified semantic encoding to obtain high-dimensional cross-modal feature embeddings and semantic representation files includes: Apply a multi-head attention mechanism to the multi-modal dataset, calculate the attention weights within and between each modality to obtain attention-weighted features; Input the attention-weighted features into a multi-layer Transformer encoder for deep feature fusion to obtain a fused feature representation; Apply a contrastive learning loss function to the fused feature representation to optimize the semantic consistency of the features and obtain semantically aligned feature vectors; Project the semantically aligned feature vectors into a high-dimensional space through a non-linear mapping layer to construct a feature relationship graph, obtaining high-dimensional cross-modal feature embeddings and semantic representation files.

[0009] Optionally, in the third implementation manner of the first aspect of the present invention, corresponding the local task requirements of multiple agents to cross-modal semantics based on the high-dimensional cross-modal feature embeddings and semantic representation files to obtain a task requirement semantic mapping, and according to the task requirement semantic mapping, performing task division and time arrangement of the multiple agents to obtain a multi-agent cooperation strategy file, including: Apply the t-SNE algorithm to the high-dimensional cross-modal feature embeddings and semantic representation files for dimensionality reduction visualization, and use the DBSCAN density clustering algorithm to cluster the dimensionality-reduced data by task type to obtain task category labels; Construct a task-semantic two-way mapping table based on the task category labels and the original task descriptions to obtain a task requirement semantic mapping, and based on the task requirement semantic mapping, analyze the forward and backward dependencies and resource requirements between tasks to construct a task dependency graph; Use a graph partitioning algorithm to partition the task dependency graph to obtain a preliminary task allocation plan, and apply a genetic algorithm to the preliminary task allocation plan to optimize the task execution order and time allocation, generating a multi-agent cooperation strategy file containing operation instructions and communication protocols for each agent.

[0010] Optionally, in the fourth implementation manner of the first aspect of the present invention, the constructing a task-semantic two-way mapping table based on the task category labels and the original task descriptions to obtain a task requirement semantic mapping, and based on the task requirement semantic mapping, analyzing the forward and backward dependencies and resource requirements between tasks to construct a task dependency graph includes: Calculate the semantic similarity between the task category labels and the original task descriptions to obtain a task-semantic association matrix; According to the task-semantic association matrix, use the bidirectional maximum matching algorithm to construct a task-semantic two-way mapping table to obtain a task requirement semantic mapping; Based on the task requirement semantic mapping, extract the temporal information and resource requirement keywords in the task description to obtain a task attribute vector; Use a graph neural network to model the task attribute vector to construct an initial task relationship graph to obtain task node representations; Apply the graph attention mechanism to the task node representation, calculate the dependence strength between tasks, and obtain the task dependence relationship graph.

[0011] Optionally, in the fifth implementation manner of the first aspect of the present invention, corresponding the local task requirements of multiple agents to cross-modal semantics based on the high-dimensional cross-modal feature embedding and semantic representation file, obtaining a task requirement semantic mapping, and performing task division and time arrangement of the multiple agents according to the task requirement semantic mapping, the obtained multi-agent cooperation strategy file includes: Use a distributed stream processing framework to perform real-time sampling and feature extraction on the key data in the execution process of multiple agents, and obtain a real-time feature sequence; Calculate the Wasserstein distance between the real-time feature sequence and a preset high-dimensional semantic representation, and apply the exponential weighted moving average method for smoothing processing to obtain a concept deviation metric value; Apply the CUSUM algorithm to the concept deviation metric value for cumulative sum analysis, set a dynamic threshold, and obtain a concept drift detection result; When the concept drift detection result exceeds the dynamic threshold, use incremental learning and knowledge distillation techniques to update the parameters of the multi-modal large language model.

[0012] Optionally, in the sixth implementation manner of the first aspect of the present invention, applying the CUSUM algorithm to the concept deviation metric value for cumulative sum analysis, setting a dynamic threshold, and obtaining a concept drift detection result includes: Perform sliding window segmentation on the concept deviation metric value to obtain time series data segments; Use the adaptive kernel density estimation method to perform probability distribution modeling on the time series data segments to obtain the current concept distribution; According to the current concept distribution, calculate the CUSUM statistic, and compare it with the historical CUSUM statistic to obtain a cumulative deviation value; Use the exponential weighted moving average algorithm to smooth the cumulative deviation value to obtain a smoothed CUSUM curve; Based on the historical data of the smoothed CUSUM curve, apply the quantile regression method to calculate the dynamic threshold to obtain an adaptive threshold function; Compare the smoothed CUSUM curve with the adaptive threshold function to determine the optimal segmentation point and obtain the concept drift detection result.

[0013] The second aspect of the present invention provides a multi-agent task cooperation device, and the multi-agent task cooperation device includes: A data acquisition module, configured to acquire visual images, text descriptions, voice instructions, and structured sensor signals, and perform spatio-temporal alignment and format conversion processing on the acquired data to obtain a multi-modal data set; A semantic encoding module, configured to perform deep fusion and unified semantic encoding on various modal inputs of the acquired multi-modal data set by using a multi-modal large language model to obtain high-dimensional cross-modal feature embeddings and semantic representation files; A task coordination module, configured to correspond the local task requirements of multiple agents with cross-modal semantics based on the high-dimensional cross-modal feature embeddings and semantic representation files to obtain a task requirement semantic mapping, and perform task division and time arrangement of the multiple agents according to the task requirement semantic mapping to obtain a multi-agent cooperation strategy file; A dynamic adjustment module, configured to, when the multiple agents execute the multi-agent cooperation strategy file, perform real-time sampling on key data and operation results during the execution process, compare the sampled data with the high-dimensional semantic representation to obtain a concept drift detection result, and dynamically adjust the multi-modal large language model when the concept drift detection result indicates a progressive drift of the concept.

[0014] A third aspect of the present invention provides a multi-agent task cooperation device, including: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor calls the instructions in the memory to enable the multi-agent task cooperation device to execute the steps of the above multi-agent task cooperation method.

[0015] A fourth aspect of the present invention provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the instructions run on a computer, the computer is enabled to execute the steps of the above multi-agent task cooperation method.

[0016] The above multi-agent task cooperation method, device, equipment, and storage medium perform deep fusion and unified semantic encoding on the acquired multi-modal data set by using a multi-modal large language model to obtain high-dimensional cross-modal feature embeddings and semantic representation files, which are used to correspond the local task requirements of multiple agents with cross-modal semantics to obtain a task requirement semantic mapping, and perform task division and time arrangement; when the multiple agents execute tasks, real-time sampling is performed on key data and operation results, compared with the high-dimensional semantic representation to obtain a concept drift detection result, and when it is detected that a concept has a progressive drift, the multi-modal large language model is dynamically adjusted. This method uses a large language model to enhance the communication and cooperation efficiency between multiple agents, optimizes task allocation and cooperation decision-making through a task requirement semantic mapping; introduces a concept drift detection and dynamic adjustment mechanism to improve long-term adaptability in a complex dynamic environment.

[0017] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention are realized and attained by the structure particularly pointed out in the specification, claims and drawings.

[0018] To make the above objectives, features and advantages of the present invention more comprehensible, the following specific preferred embodiments are given, in conjunction with the accompanying drawings, and are described in detail as follows. Description of the Drawings

[0019] Figure 1 It is a schematic diagram of the first embodiment of the multi-agent task collaboration method in the embodiments of the present invention; Figure 2 It is a schematic diagram of an embodiment of the multi-agent task collaboration device in the embodiments of the present invention; Figure 3 It is a schematic diagram of an embodiment of the multi-agent task collaboration device in the embodiments of the present invention. Detailed Embodiments

[0020] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0021] As used in the embodiments of the present invention, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include other unlisted steps or units, or may optionally further include other steps or units inherent to these processes, methods, products or devices.

[0022] For ease of understanding of this embodiment, a multi-agent task collaboration method disclosed in the embodiments of the present invention will be introduced in detail first. As Figure 1 shown, the method includes the following steps: 101. Collect visual images, text descriptions, voice commands, and structured sensor signals, and perform spatio-temporal alignment and format conversion processing on the collected data to obtain a multi-modal data set; In one embodiment of the present invention, the acquisition of visual images, text descriptions, voice commands, and structured sensor signals, and the spatio-temporal alignment and format conversion processing of the acquired data to obtain a multi-modal data set include: performing convolutional neural network processing on the visual image to extract an image feature vector; applying the BERT model to the text description for word embedding to obtain a text feature vector; extracting audio features through MFCC transformation on the voice command to obtain a voice feature vector; performing wavelet transformation on the structured sensor signal to obtain a sensor feature vector; aligning the image feature vector, text feature vector, voice feature vector, and sensor feature vector according to the time stamp to construct a multi-modal time series matrix, and applying the dynamic time warping algorithm to the multi-modal time series matrix to obtain spatio-temporally aligned data; mapping the data of different modalities in the spatio-temporally aligned data to a unified numerical range, and reducing the feature dimension through principal component analysis to obtain a feature matrix after dimensionality reduction; applying a tensor decomposition algorithm to the feature matrix after dimensionality reduction to extract shared and independent modal features, and integrating them into a tensor structure in a unified format to obtain the multi-modal data set.

[0023] Specifically, when performing convolutional neural network processing on a visual image, the original image is first input into a normalization module. By performing linear or non-linear transformation on the pixel intensity, the influence of illumination and resolution differences on feature representation is reduced. Then, a number of learnable convolutional kernels are set in the convolutional layer. These convolutional kernels scan the image layer by layer under the sliding window mechanism and calculate the local weighted sum at each position. Through this convolutional operation, low-level features such as edges or corners can be extracted in the early stage, and more abstract patterns can be captured in deeper convolutional stacks. Thereafter, the pooling layer performs downsampling in the spatial dimension to reduce redundancy and retain the most representative activation responses. In some implementations, residual connections are used to alleviate the vanishing gradient problem that occurs during network training and help the deep structure more effectively fuse features within different receptive field ranges. The finally output image feature vector is obtained after calculation in the fully connected layer or the global average pooling layer. The vector dimension is usually related to the network depth and the number of convolutional kernels. This representation helps to focus on the differences in the target texture or shape distribution in the subsequent stage. The image feature vector is stored in a specific cache and retained together with the corresponding time stamp and image metadata, providing a queryable index for subsequent alignment with text, voice, and sensor information.

[0024] Specifically, when applying the BERT model to the text description for word embedding processing, the input text is first tokenized, stop words are removed, and lemmatization is performed. Then, the serialized result is input into a multi-layer Transformer encoder. At this time, the self-attention mechanism assigns attention weights to each word within the sentence and makes full use of the multi-head attention units to capture the dependency structure between different positions in the sentence, thereby generating a text feature vector that is closely related to the context. The voice command obtains the spectral features of the audio signal through MFCC transformation. The specific steps include operations such as pre-emphasis filtering, framing, windowing, and discrete cosine transform, and a set of cepstral coefficients will be output to represent the main component distribution of the voice. Next, the structured sensor signal will be decomposed into different scales and frequency bands in the wavelet transform stage, used to disassemble the transient and periodic patterns in the original signal, and extract feature vectors that can identify the device operating state or environmental conditions. The above text feature vector, voice feature vector, and sensor feature vector are stored together with the image feature vector through a unified index management module, and source identifiers and time stamps are marked at the data structure level to provide a multi-modal data synchronization record for subsequent alignment.

[0025] Specifically, the image feature vector, text feature vector, voice feature vector, and sensor feature vector are aligned in an orderly manner based on the acquisition time, and the dynamic time warping algorithm is applied to the constructed multi-modal time series matrix. This matrix uses the time dimension as the horizontal axis and is used to store multi-modal feature values collected at the same or similar times. There will be a time difference between different modalities due to hardware sampling rates or communication delays. The dynamic time warping algorithm uses the method of elastic matching to find the optimal path to make the feature curves consistent on the time axis. After alignment, spatio-temporal alignment data will be output, which is used to accurately map the same phenomenon or event recorded by each modality to the same or adjacent time windows. In this process, correction means such as interpolation or filtering will be taken for missing values or abnormal samples to ensure the continuity of the overall data stream. The aligned results will be blocked or labeled according to the feature type, and retrievable index fields will be retained in the time series matrix so that subsequent algorithms can efficiently reference and associate these aligned multi-modal information.

[0026] Specifically, map the different modal features in the above spatio-temporal aligned data to a unified numerical range, and perform dimensionality reduction processing using principal component analysis. The numerical range mapping is generally achieved through Z-score or min-max normalization methods to avoid uneven feature distribution caused by differences in dimensions among modalities. Principal component analysis calculates the covariance within a high-dimensional matrix and performs eigen decomposition on it, so as to select the first several principal components with a relatively high cumulative contribution rate to construct the feature matrix after dimensionality reduction. Subsequently, in the tensor decomposition algorithm, the dimensionality reduction matrix will be regarded as part of a high-order tensor, and shared features and independent features are obtained through decomposition and reconstruction. The shared features represent the intersection of multi-modal data in the latent structure, while the independent features correspond to the information specific to a particular modality. After the tensor decomposition process, all results are combined into a unified tensor structure to complete cross-modal parallel analysis and correlation retrieval using a single data format in subsequent steps. The finally generated multi-modal dataset has a relatively compact representation form and the ability to focus on key features, accommodating the internal connections and differences of different modal information in the same coordinate space.

[0027] 102. Use a multi-modal large language model to deeply fuse and uniformly semantically encode various modal inputs of the collected multi-modal dataset to obtain high-dimensional cross-modal feature embeddings and semantic representation files; In an embodiment of the present invention, the use of a multi-modal large language model to deeply fuse and uniformly semantically encode various modal inputs of the collected multi-modal dataset to obtain high-dimensional cross-modal feature embeddings and semantic representation files includes: applying a multi-head attention mechanism to the multi-modal dataset, calculating the attention weights within and between each modality to obtain attention-weighted features; inputting the attention-weighted features into a multi-layer Transformer encoder for deep feature fusion to obtain a fused feature representation; applying a contrastive learning loss function to the fused feature representation to optimize the semantic consistency of the features to obtain semantically aligned feature vectors; projecting the semantically aligned feature vectors into a high-dimensional space through a non-linear mapping layer to construct a feature relationship graph to obtain high-dimensional cross-modal feature embeddings and semantic representation files.

[0028] Specifically, when applying the multi - head attention mechanism to the multi - modal dataset, it is necessary to first load the vector representations of different modalities such as images, texts, voices, and sensors into the corresponding attention calculation modules, and assign trainable query, key, and value matrices to each modality. The multi - head attention mechanism will perform multiple groups of attention calculations in parallel to capture the correlation degrees within and between modalities. In specific implementation, the attention scores are obtained through vector dot - product, and after scaling and Softmax normalization, a set of weight coefficients are generated. These coefficients will be multiplied with the corresponding keys and values to highlight the regions or features most relevant to the current decision. The attention allocation within a modality can strengthen the feature differences at different positions in the same modality, and the attention allocation between modalities can highlight or suppress certain heterogeneous features during cross - modal information fusion. Then, the obtained weighted results are concatenated to form attention - weighted features, which are used to retain the complementary information of each modality in the spatio - temporal dimension while weakening the excessive influence in terms of dimension or noise. The output result is presented in the form of a tensor and retains a traceable index for subsequent reference.

[0029] Specifically, the above - mentioned attention - weighted features are input into a multi - layer Transformer encoder for in - depth feature fusion. In the specific process, the encoder contains several stacked sub - layers, and each sub - layer contains a multi - head attention sub - layer and a feed - forward network sub - layer. The multi - head attention sub - layer performs a re - association calculation on the weighted features of the previous step, and the feed - forward network sub - layer adjusts the distribution of features through linear mapping and activation functions. Residual connections can establish a direct path between the input and output, allowing the gradient to propagate smoothly during training and making it easier for the encoder to integrate inter - layer information. Layer normalization corrects the mean and variance of the output of each sub - layer to avoid numerical explosion or decay in some dimensions. After multiple layers of stacking, the fused feature representation retains the relatively complete temporal and spatial correlations in the cross - modal context, and can uniformly describe the differences and connections of multiple modalities in the same vector space.

[0030] Specifically, applying a contrastive learning loss function to the fused feature representation aims to optimize the intrinsic consistency of cross - modal features at the semantic level. In implementation, first construct the index relationships of positive and negative sample pairs. Positive sample pairs refer to modality combinations that are similar in semantics or labels, while negative sample pairs are modality combinations with large semantic differences. Contrastive learning measures the feature distances between positive sample pairs and between negative sample pairs, and requires the former to be reduced and the latter to be increased in the loss function. Through backpropagation and gradient update, the ability of the model to distinguish similar and dissimilar concepts will be gradually improved. This process is iterated through multiple training batches, and different positive and negative sample pairs of different modalities are dynamically sampled in each batch. When the contrastive learning is completed, the output semantic alignment feature vectors are more structured in the geometric space, and can more stably reflect the same or similar entities expressed by different modalities, thus ensuring the accuracy of semantic understanding.

[0031] Specifically, the semantically aligned feature vector is projected into a high-dimensional space through a nonlinear mapping layer to construct a feature relationship graph. In implementation, a multi-layer fully connected network or MLP (multi-layer perceptron) is set during projection, and the model's fitting ability for complex data distribution is improved through activation functions, and a high-dimensional vector is obtained at the network output. The feature relationship graph can store high-dimensional vector nodes and their similarities or attention connections with the help of a graph structure, and the edge weights of the graph can be calculated based on cosine similarity or Euclidean distance. In this graph structure, similar cross-modal vectors are aggregated into local clusters, and vectors with large differences are separated from each other, thereby forming a visual or retrievable cross-modal partition. The high-dimensional cross-modal feature embedding and semantic representation file can be used for subsequent task assignment, semantic reasoning, or strategy generation, and the potential modal association path can be traced back by retrieving the relationship between graph nodes during query.

[0032] 103. Based on high-dimensional cross-modal feature embedding and semantic representation files, the local task requirements of multiple agents are matched with cross-modal semantics to obtain task requirement semantic mapping. Based on the task requirement semantic mapping, the task division and time arrangement of multiple agents are carried out to obtain the multi-agent collaborative strategy file; In one embodiment of the present invention, based on the high-dimensional cross-modal feature embedding and semantic representation file, the local task requirements of multiple agents are matched with cross-modal semantics to obtain a task requirement semantic mapping, and according to the task requirement semantic mapping, the task division and time arrangement of the multiple agents are performed to obtain a multi-agent collaborative strategy file, including: applying the t-SNE algorithm to the high-dimensional cross-modal feature embedding and semantic representation file for dimensionality reduction visualization, and using the DBSCAN density clustering algorithm to cluster the task types of the reduced-dimensional data to obtain task category labels; constructing a task-semantics bidirectional mapping table based on the task category labels and the original task description to obtain a task requirement semantic mapping, and based on the task requirement semantic mapping, analyzing the forward and backward dependencies and resource requirements between tasks to construct a task dependency graph; using a graph segmentation algorithm to divide the task dependency graph to obtain a preliminary task allocation plan, and applying a genetic algorithm to the preliminary task allocation plan to optimize the task execution order and time allocation, and generating a multi-agent collaborative strategy file containing the operation instructions and communication protocols of each agent.

[0033] Specifically, on the premise that the high-dimensional cross-modal feature embedding and semantic representation files have been preprocessed and stored in a specified data structure, it is necessary to use the t-SNE algorithm to perform dimensionality reduction and visualization operations on the vectors therein. When implementing, first, all the embedding vectors are packed into a matrix, and the low-dimensional coordinates are initialized randomly. Subsequently, t-SNE calculates the similarity probability distribution between any two points in the high-dimensional space and iteratively updates the coordinate positions in the low-dimensional space to keep the KL divergence between the two distributions low, thereby reproducing the neighborhood structure among the original data in a two-dimensional or three-dimensional plane. To reduce the interference of local extrema and accelerate the convergence process, t-SNE often uses a higher learning rate in the early stage and dynamically adjusts the step size in the later stage, and monitors the projection quality according to the cost function after each iteration. After dimensionality reduction, a coordinate table of all samples in the low-dimensional coordinate system is obtained for visualization or subsequent clustering. Next, the DBSCAN density clustering algorithm is introduced to read this dimensionality reduction result and cluster the point cloud. This algorithm calculates the number of neighbors for each point and compares it with the core radius threshold to identify high-density connected regions. If a point does not meet the core point standard within the neighborhood range, it will be marked as noise or a boundary. This can effectively distinguish the data distribution in an unsupervised scenario. The output result includes a set of clustering label files, where each cluster corresponds to a certain degree of semantic consistency or functional similarity, facilitating the determination of task types based on these labels in subsequent steps. This label file will be saved in a unified structure, and the labels are mapped to the corresponding point coordinates and vector indices, providing an accurate basis for subsequent matching steps.

[0034] Specifically, based on the task category labels output by DBSCAN and the original task description text, a task-semantic bidirectional mapping table needs to be constructed to correlate the clustering results with the original requirements. First, perform similarity measurement at the word vector or sentence vector level on the task category labels and text descriptions, and calculate the matching degree through cosine similarity or Euclidean distance. If there is a high semantic similarity between the key topic words of a category label and a certain task description, a two-way link will be generated in the mapping table. This step usually unifies the text tokenization method and the dictionary first to ensure that there are no tokenization conflicts or deviations during matching. To further ensure the accuracy of the corresponding relationship, a bidirectional maximum matching algorithm is adopted: on the one hand, search for the most matching task description starting from the label, on the other hand, retrieve the most matching label starting from the task description, and then synthesize these two results to determine the optimal mapping. After obtaining a bidirectional mapping table with a high degree of completion, the resource requirements and dependencies between tasks will be analyzed in depth. The resource requirements analysis will retrieve the references to hardware, network, personnel, or environmental conditions in the original description and cross-modal semantics, and extract the shared elements and exclusive elements. The dependencies are marked with directed edges based on the temporal constraints, condition triggers, or domain logic constraints in the task description, and all nodes (tasks) are stored in a graph structure. The resulting task dependency graph will show which subtasks need to be completed first or under what conditions to enter the next process when multiple agents execute these tasks, and provide a more intuitive computable representation for the subsequent scheduling algorithm.

[0035] Specifically, based on the task dependency graph and the task-semantic bidirectional mapping table, it is necessary to complete the preliminary task allocation through a graph partitioning algorithm and further use a genetic algorithm to optimize the execution order and time arrangement. The graph partitioning algorithm splits the subgraphs according to the node correlation or edge weight. Usually, objective functions such as minimum cut or modularity maximization are selected to group nodes with close internal connections in the same subgraph, thereby reducing the coupling between subgraphs. After the partitioning is completed, each subgraph in the preliminary allocation plan can be regarded as a local division unit, and the internal nodes point to multiple related tasks. However, there may be redundancy or timing conflicts in this allocation, and a genetic algorithm is needed to globally optimize the sorting and time window. In the specific implementation, the task order of each subgraph is first encoded into a chromosome sequence, and the fitness function measures the total execution duration, resource occupancy conflict, and degree of dependency violation. Then, new populations are generated through crossover and mutation operations. At the end of each generation of iteration, the individual with the highest fitness is selected and retained, and the individuals with low fitness are replaced. As the number of evolutionary generations increases, the rationality of the task scheduling gradually improves until the maximum number of iterations or the convergence criterion is reached. The final output multi-agent collaboration strategy file contains the task list that each agent should execute, the detailed time arrangement, and the communication protocol between agents. This file is usually presented in a structured or scripted form and can be directly called and started in a multi-node scenario to initiate the multi-agent collaboration process, thereby forming a more efficient task management and execution mode.

[0036] Further, constructing the task-semantic bidirectional mapping table according to the task category label and the original task description, obtaining the task requirement semantic mapping, and based on the task requirement semantic mapping, analyzing the front and back dependency relationships and resource requirements between tasks, the steps of constructing the task dependency graph include: calculating the semantic similarity between the task category label and the original task description to obtain the task-semantic association matrix; according to the task-semantic association matrix, using the bidirectional maximum matching algorithm to construct the task-semantic bidirectional mapping table to obtain the task requirement semantic mapping; based on the task requirement semantic mapping, extracting the timing information and resource requirement keywords in the task description to obtain the task attribute vector; using a graph neural network to model the task attribute vector to construct the initial task relationship graph to obtain the task node representation; applying the graph attention mechanism to the task node representation to calculate the dependency strength between tasks to obtain the task dependency graph.

[0037] Specifically, it is necessary to calculate the semantic similarity between the task category labels and the original task descriptions to obtain a task-semantic association matrix. During implementation, all labels and descriptions can be first converted into comparable text vector representations, usually achieved through word segmentation, lemmatization, and word vector models. Subsequently, the labels and descriptions are paired one by one according to the same vector dimension to calculate the similarity. The similarity measure can choose cosine similarity or Euclidean distance. Cosine similarity takes the ratio of the dot product and norm of two vectors as the similarity score, while Euclidean distance directly measures their geometric distance in the vector space. After the calculation, the similarity values between each label and each description are filled into a two-dimensional matrix. The row index is used to identify the label, and the column index is used to identify the description. At this time, each element in the matrix corresponds to the matching degree of a pair of "label-description", and the higher the value, the closer their semantic association. If dealing with ultra-large-scale data, the labels and descriptions can be chunked and sliced in a parallel framework, and the similarity calculation can be executed collaboratively on multiple computing nodes to accelerate the matrix generation process. After generation, the task-semantic association matrix is saved in a unified format. A common method is to store the matrix in a database or a memory-mapped file, and retain the index and numerical precision description to facilitate accurate positioning of the corresponding similarity scores when retrieving labels or descriptions in subsequent steps. Under the action of the two-way indexing of rows and columns, this matrix provides a directly referenceable scoring basis for the subsequent bidirectional maximum matching algorithm and also provides a necessary data mapping path for constructing task dependencies later.

[0038] Specifically, based on the task-semantic association matrix generated in the previous step, it is necessary to use the bidirectional maximum matching algorithm to construct a task-semantic two-way mapping table, thereby forming a clear semantic mapping of task requirements. During implementation, the description with the highest similarity for each label can be first retrieved in the matrix, and this pair of "label-description" is temporarily stored as a candidate match. Subsequently, the most suitable label is found for each description. If this description has been selected by another label before and there are conflicts in terms of similarity or priority between the two, it is necessary to compare between the conflicting parties to determine which matching combination is more in line with the maximum matching principle. The maximum matching algorithm will continuously iterate and update between rows and columns, gradually eliminating duplicate pairings or local conflicts until reaching the global optimum or unable to improve the matching score any further. At the implementation level, the matrix can be processed through the Hungarian algorithm or similar linear programming methods, or heuristic methods can be used to quickly approximate the optimal solution in large-scale scenarios. After the matching is completed, a mapping table will be generated, where each record indicates which task description a specific label corresponds to and the corresponding semantic score. This mapping table is presented in a two-way manner: it can be indexed from the label to the description, or from the description to the label, facilitating the retrieval or modification of this association relationship in subsequent links. Once generated, it forms a semantic mapping of task requirements to identify the optimal semantic matching between task categories and text descriptions.

[0039] Specifically, on the premise of obtaining the semantic mapping of task requirements, it is necessary to parse the corresponding text description, extract the timing information and resource requirement keywords, and finally construct a task attribute vector. In specific implementation, first find the description text corresponding to each tag from the mapping table through indexing, and then perform word segmentation and part-of-speech recognition in the natural language processing module, and label the semantic attributes of each word (such as time markers, device requirements, or condition trigger words, etc.). For expressions with an operation sequence or time window, such as "B can only be executed after A is completed" or "Device C needs to be continuously monitored for 10 minutes", its timing attributes will be extracted separately and transcoded into numerical features, such as start time, duration, and dependency identifier. For resource requirement keywords, information such as device names, network bandwidth requests, or manpower allocation restrictions that may appear will be mined at the sentence level or phrase level, and corresponding fields will be set in the attribute vector to record this requirement. If a description contains multiple requirement elements, they will be retained in multiple dimensions in the attribute vector to avoid information loss. After text parsing and numericalization, the formed task attribute vector is usually stored as a fixed-length vector or a sparse vector, and each vector corresponds one-to-one with the corresponding tag (i.e., task type), and its source description is specified in the index structure, so that relevant timing and resource attributes can be quickly retrieved during subsequent modeling and graph structure generation.

[0040] Specifically, after extracting the task attribute vector, it is necessary to use a graph neural network to model it into an initial task relationship graph to obtain the task node representation. First, create a node object for each task tag and write the corresponding attribute vector into the node features. Next, build an initial edge structure by detecting potential similarity or conflict information between the attribute vectors. For example, if two tasks share some parts in resource requirements or have a sequential relationship in timing logic, an edge can be generated between them, and the relationship type or conflict intensity is written in the attribute of the edge. Then, input this graph into a graph neural network, usually using a Graph Convolutional Network, a Graph Attention Network, or other models that can handle node features and edge relationships. During training, the graph neural network samples the node neighborhood in batches, iteratively aggregates the attribute vector and neighbor node information, and gradually updates the node embedding representation under the action of multi-layer convolution or attention mechanism. After each round of iteration, a certain graph-level or node-level loss function is calculated to ensure that the node representation can not only represent its own attributes but also reflect the local graph structure. If there are labels or constraints, they can be incorporated into semi-supervised or fully supervised training to make the node representation more distinguishable. After training, the node representation will be exported to a mapping table, indicating the correspondence between each task node and its high-dimensional vector features, laying an operable numerical foundation for the subsequent calculation of dependency strength.

[0041] Specifically, after obtaining the initial task relationship graph and the corresponding node representations, it is necessary to use the graph attention mechanism to calculate the dependence strength between tasks and generate the task dependence relationship graph. In specific implementation, each edge is traversed in the graph structure in turn. For the two nodes connected by the edge, their embedding vectors are read respectively, and the relevance between the nodes is calculated through a trainable attention weight matrix or affine transformation. Then, the Softmax operation is performed on the relevance of all neighbors of the same node, and the result is normalized to the attention weight distribution. The larger the weight, the higher the importance of the neighbor to the current node. To more accurately reflect the influence of the time series and resource dimensions in the dependence relationship graph, the time series and resource fields obtained in the third and fourth steps can be introduced into the input of the attention calculation, and weighted implementation is achieved in the attention scoring by combining these attributes. Finally, after obtaining a set of attention weights on all edges, the numerical values can be regarded as the dependence strength between tasks, and the numerical values are stored in the edge attributes of the graph. If the attention weight exceeds a certain threshold or conforms to the preset logic, the dependence direction or constraint conditions of the tight coupling are marked in the task dependence relationship graph. The task dependence relationship graph obtained in this way not only presents the topological connection of each node, but also quantifies the dependence strength in a numerical way, which is convenient to arrange the task order or resource allocation according to these strengths in the subsequent scheduling or optimization algorithms, and provides a visual and computable relationship basis for large-scale multi-agent cooperation.

[0042] 104. When multi-agents execute the multi-agent cooperation policy file, key data and operation results during the execution process are sampled in real time, and the sampled data is compared with the high-dimensional semantic representation to obtain the concept drift detection result. When the concept drift detection result shows a progressive drift of the concept, the multi-modal large language model is dynamically adjusted.

[0043] In an embodiment of the present invention, corresponding the local task requirements of multi-agents with the cross-modal semantics based on the high-dimensional cross-modal feature embedding and the semantic representation file to obtain the task requirement semantic mapping, and performing the task division and time arrangement of the multi-agents according to the task requirement semantic mapping to obtain the multi-agent cooperation policy file includes: using a distributed stream processing framework to sample and extract features of key data during the execution process of multi-agents in real time to obtain a real-time feature sequence; calculating the Wasserstein distance between the real-time feature sequence and the preset high-dimensional semantic representation, and performing smoothing processing by applying the exponentially weighted moving average method to obtain the concept drift metric value; applying the CUSUM algorithm to the concept drift metric value for cumulative sum analysis and setting a dynamic threshold to obtain the concept drift detection result; when the concept drift detection result exceeds the dynamic threshold, updating the parameters of the multi-modal large language model by using incremental learning and knowledge distillation techniques.

[0044] Specifically, when using a distributed stream processing framework to perform real-time sampling and feature extraction on the key data of the multi-agent execution process, it is necessary to deploy lightweight data collection agents on each agent node and set up a distributed message queue or stream processing module on the central or edge server side. During the execution process, each node will generate various forms of observation data, such as visual recognition results, robotic arm position coordinates, voice interaction records, etc. These data will be packed by the collection agent with timestamps attached and then transmitted to the stream processing engine. When the stream processing engine receives the incoming data stream, it will segment and summarize the data according to the predefined window policy and perform cleaning or simple feature statistics operations in the specified operator. If more fine-grained feature extraction is required, a pre-set convolutional network or time series analysis model can be connected inside the operator to perform embedded encoding on video frames or sensor sequences and then output shorter and structured feature vectors. For voice input, it will first go through VAD endpoint detection, and then be filtered, framed, and Mel cepstral coefficients will be extracted to obtain comparable audio features. Each feature stream will perform parallel computing in a distributed environment to reduce the bottleneck brought by centralized processing, and after merging downstream, form a multi-dimensional real-time feature sequence. This sequence is arranged in chronological order, and each moment contains the feature data of the current task state or execution scene, which helps to continuously monitor the overall dynamics and local behaviors of multi-agents under a large throughput and provides an accurate data carrier for subsequent offset measurement or concept correction.

[0045] Specifically, when calculating the Wasserstein distance between the above real-time feature sequence and the preset high-dimensional semantic representation, it is necessary to first normalize the multi-dimensional vectors extracted from each time window in the real-time sequence to ensure that the numerical distribution is in the same dimension as the preset semantic representation. Subsequently, a Wasserstein distance metric based on the optimal transport theory can be selected to measure the difference between the real-time feature distribution and the high-dimensional semantic vector distribution. In specific implementation, the real-time vector set will be regarded as the probability distribution P, and the semantic representation will be regarded as another distribution Q. By solving the optimal transport plan, the cumulative cost corresponding to the set of transport schemes that minimize the cost among all feasible mappings will be calculated. Compared with the simple Euclidean distance, the Wasserstein distance pays more attention to the distribution shape and overall matching degree, so it can more accurately reflect the deviation degree between the task execution state and the semantic representation in a multi-modal scenario. After the calculation, a set of distance values that evolve over time will be obtained. To suppress short-term jitter and highlight trend features, an exponentially weighted moving average method can be applied to this set of distance values. This method superimposes the decreasing weights of historical values on the current distance value, thereby smoothing high-frequency noise and forming a concept offset metric value. This metric value is continuously stored in time series and can indicate the stability or offset trend of task execution in the current semantic space.

[0046] Specifically, after obtaining the concept offset metric value, it is necessary to apply the CUSUM algorithm to this metric value for cumulative sum analysis and set a dynamic threshold to determine whether a concept drift phenomenon occurs. When implementing, each metric value will be processed in chronological order, and a CUSUM-based incremental statistic S will be established. The update rule of this statistic is usually \(S_t=\max(0, S_{t - 1}+(\text{offset metric value}-\text{reference mean}-\text{tolerance interval}))\). If \(S_t\) is always greater than zero and gradually increases after multiple iterations, it indicates that the gap between the metric value and the reference baseline is accumulating. In order to automatically adapt to environmental changes at different stages, a dynamic threshold needs to be set through an adaptive strategy. This threshold can be updated in real time according to quantile regression or the distribution of historical offset data and compared with the current \(S_t\). When \(S_t\) exceeds this threshold, it can be judged that the concept drift trend is relatively obvious. The CUSUM algorithm has good sensitivity to both sudden deviations and slow deviations during the cumulative increment process, and can provide timely detection indications for the common gradual concept drift or abnormal offset in large-scale multi-agent collaboration, and mark the time and intensity level of each offset occurrence point numerically on the time series.

[0047] Specifically, when the concept drift detection result exceeds the dynamic threshold, the parameters of the multi-modal large language model are updated using incremental learning and knowledge distillation techniques to locally correct the model without interrupting the overall collaboration task. When implementing, first extract the samples with abnormal offsets that have recently occurred from the historical buffer or the latest samples of the distributed stream processing framework and mix them with a batch of previously normal data. Incremental learning will load this part of the new data on the basis of the original weights of the model, and iteratively update the parameters such as the attention mechanism, projection layer, or loss function inside the model through mini-batch training or online optimization to capture changes in the environment or tasks. To avoid overfitting of the new data or forgetting old knowledge, knowledge distillation techniques can be deployed during the update process, and the output distribution of the new model is constrained by the old model or a larger-scale teacher model, so that the new model can not only absorb the correction information of recent offsets but also maintain the original cross-modal understanding ability. After the incremental learning is completed, the updated model weights and concept mapping tables will be synchronized to all agents or edge nodes that need to collaborate in the distributed environment to maintain global semantic consistency and prevent concept drift from accumulating again in the subsequent execution stage. The final completed model update record can be retained in the system log, including the update time, sample distribution characteristics, and parameter change range, which is used to provide a traceable basis for later traceability or re-tuning.

[0048] Further, applying the CUSUM algorithm to the concept offset metric value for cumulative summation analysis and setting a dynamic threshold, the concept drift detection result is obtained as follows: segmenting the concept offset metric value by a sliding window to obtain time series data segments; using the adaptive kernel density estimation method to model the probability distribution of the time series data segments to obtain the current concept distribution; calculating the CUSUM statistic according to the current concept distribution and comparing it with the historical CUSUM statistic to obtain the cumulative deviation value; using the exponential weighted moving average algorithm to smooth the cumulative deviation value to obtain a smoothed CUSUM curve; applying the quantile regression method to calculate the dynamic threshold based on the historical data of the smoothed CUSUM curve to obtain an adaptive threshold function; comparing the smoothed CUSUM curve with the adaptive threshold function to determine the optimal segmentation point and obtain the concept drift detection result.

[0049] Specifically, when segmenting the concept offset metric value by a sliding window to obtain time series data segments, one or more time windows with fixed or adaptive lengths need to be defined on the continuously sampled concept offset metric data, and the metric values are encapsulated into small-scale segments within each window. In implementation, the window size is first determined according to the system refresh frequency or the update period of the business scenario. By traversing the entire concept offset metric sequence, adjacent data points are assigned to the corresponding windows to ensure that each window contains relatively complete and continuous metric information. If data loss or time misalignment occurs during the execution, interpolation or extrapolation methods can be used for smoothing correction to avoid generating excessive noise in subsequent analysis. After segmenting the sequence data, an ordered list of time series data segments is obtained, where each element represents the concept offset state within a certain time range. To track the law of concept evolution over time, these data segments retain their start and end times and statistical features in the index structure, such as the average value, variance, or frequency domain features within the window, to provide necessary context references for subsequent algorithms during modeling. In some cases, the overlapping window technique is also introduced according to specific requirements to achieve partial data sharing between adjacent windows, enhance the ability to capture sudden anomalies, and avoid missing key signals at the window boundaries. The finally output time series data segments not only have the index retrieval function but can also be compared with other monitoring metrics for comprehensively evaluating the stability of real-time task execution and the consistency of concept understanding.

[0050] Specifically, when using the adaptive kernel density estimation method to model the probability distribution of the time series data segment and obtain the current concept distribution, it is necessary to first regard the concept offset values within each data segment as a set of independent observation samples, and find their true distribution shape through kernel density estimation without any assumptions. In implementation, an initial bandwidth can be set first and a suitable kernel function can be selected, such as a Gaussian kernel or a dual-core mixed function, and then the kernel distance from each observation point to all sample points is calculated to obtain the density estimation value. If there is a large variance or multimodal distribution in the concept offset values within the data segment, in order to make the estimation result more refined, an adaptive bandwidth strategy needs to be adopted, that is, different bandwidth sizes are used for regions with relatively high or low local density, reducing the bandwidth in the dense region to improve the resolution and increasing the bandwidth in the sparse region to avoid the estimation curve being too steep. After the adaptive kernel density estimation is completed, one or more smooth probability distribution curves will be output, which are used to characterize the concentration degree and diffusion degree of the concept offset measurement values in the numerical space under the current window. If kernel density estimation is performed on multiple time window segments, their distribution functions can be stacked or continuously plotted on the time axis to observe the evolution pattern of the concept in the long time series. The finally obtained current concept distribution will be used in the subsequent steps to calculate the CUSUM statistic, and when comparing with the historical distribution, it can highlight the cumulative effect or phased change of the concept offset.

[0051] Specifically, according to the current concept distribution, calculate the CUSUM statistic and compare it with the historical CUSUM statistic to obtain the cumulative deviation value. It is necessary to estimate the degree of deviation from the reference mean at each time window or data segment and continuously accumulate the deviation signal. In implementation, a reference baseline will be selected, such as the average level of the concept offset measurement value in the stable stage, and then the distribution information within each data segment will be integrated, and the difference between the center or eigenvalue of the distribution and the reference baseline will be used as the incremental signal for this segment. The CUSUM algorithm will maintain a rolling statistic S_t, usually accumulated according to S_t = S_{t - 1} + (increment - control parameter) or other forms. If S_t shows a gradual increase after multiple iterations, it indicates that the concept offset has occurred repeatedly in multiple time periods. In order to identify the directional offset, two independent CUSUM sequences, positive and negative, can be set to distinguish whether the concept has an upward or downward deviation. When the CUSUM update for each data segment is completed, the CUSUM statistic at this moment needs to be differentiated or directly compared with the CUSUM statistic at the historical moment to obtain the cumulative deviation value. If the cumulative deviation value shows an obvious increasing trend, it indicates that there is a certain persistence in the concept offset, otherwise it may indicate that the concept offset belongs to short-term fluctuations or system noise. This continuous tracking of the time series provides a basis for accumulating segment by segment for judging the dynamic threshold later, and also enables the entire detection mechanism to still have the ability to sensitively capture slow or intermittent concept changes.

[0052] Specifically, when using the exponentially weighted moving average algorithm to smooth the cumulative deviation value and obtain the smoothed CUSUM curve, it is necessary to first collect the CUSUM statistics output in previous iterations and regard them as a sequence that changes over time. During implementation, the weighted synthesis of each latest value and the previous historical value of this sequence will be performed, and a higher weight coefficient will be given to more recent data to emphasize the impact of the latest trend on the overall curve. The weighting process often uses the formula EWMA_t = α * current value + (1 - α) * EWMA_{t-1}, where α represents the smoothing coefficient, which is set by the user or the system according to the characteristics of the scenario. In this way, the large fluctuations caused by instantaneous spikes or random fluctuations on the curve can be suppressed, making the change trend of the CUSUM statistics easier to observe and measure. After the processing is completed, a smoothed CUSUM curve will be generated, and each point on the curve represents the cumulative and weighted result of the previous deviation value at time t. If the concept deviation changes frequently during a certain period, the curve will smoothly show a steeper rising or falling shape; if the concept state is relatively stable, the curve will fluctuate within a certain range and tend to be smooth.

[0053] Specifically, when calculating the dynamic threshold using the quantile regression method based on the historical data of the smoothed CUSUM curve to obtain the adaptive threshold function, it is necessary to collect a certain number of historical smoothed CUSUM curve values and perform statistical analysis on them along the time axis or in batch samples. Quantile regression will fit a function for a specified quantile (such as 0.95 or 0.99, etc.) to delimit the high-risk interval in each time period or feature domain. When the fluctuation amplitude of the CUSUM curve breaks through the regression line corresponding to this quantile, it can be regarded as a serious concept deviation or a significant cumulative trend. The implementation process can first select an appropriate quantile according to the distribution characteristics of the historical curve and use a linear or non-linear regression model to fit the shape of the curve. After fitting, one or more quantile functions will be output to provide variable thresholds in different time ranges. If the curve fluctuates generally greatly during a certain period, the regression function will automatically raise the threshold to avoid excessive misjudgment of normal fluctuations; if the curve is overall stable, the threshold will decrease accordingly to enhance the sensitivity to minor deviations. Finally, the adaptive threshold function will be combined with the time index in the data structure and provided externally in a queryable manner. When the external detection module needs to determine whether the CUSUM curve exceeds the boundary, it can directly call this function to obtain the threshold at the corresponding moment.

[0054] Specifically, when comparing the smoothed CUSUM curve with the adaptive threshold function to determine the optimal segmentation point and obtain the concept drift detection result, it is necessary to traverse the values of the CUSUM curve at each time point and subtract or compare them with the output of the threshold function corresponding to the same time point. If the curve in a certain time period exceeds the threshold and maintains a certain duration or cumulative number of times, this can be determined as the concept drift section. In the implementation process, the difference between the curve and the threshold function is used as the judgment basis. If the difference is positive and continuously greater than zero, it indicates that the concept deviation accumulation has exceeded the normal range; if this difference is still in the negative value range or alternates frequently, it can be regarded as a temporary oscillation or ordinary fluctuation, and continue to track to exclude false alarms caused by excessive sensitivity. When it is determined that a concept drift occurs, a segmentation point or marked interval is recorded on the curve to indicate that the concept understanding or representation of the system has been significantly different from before after this time. If more stringent detection is required, it can be required that the difference amplitude exceeds a certain order of magnitude or shows abnormalities in multiple consecutive windows to prevent inaccurate judgments caused by short-term spikes. After the detection is completed, a report including the start and end times of the drift interval, the cumulative offset intensity, and the anomaly level is output, providing a quantifiable reference for subsequent incremental learning or model parameter adjustment.

[0055] In this embodiment, a multi-modal large language model is used to deeply fuse and uniformly semantically encode the collected multi-modal data set to obtain high-dimensional cross-modal feature embeddings and semantic representation files, which are used to correspond the local task requirements of multi-agents with cross-modal semantics to obtain task requirement semantic mappings for task division and time arrangement; when multi-agents execute tasks, key data and operation results are sampled in real time and compared with the high-dimensional semantic representation to obtain the concept deviation detection result. When it is detected that a concept shows a progressive drift, the multi-modal large language model is dynamically adjusted. This method uses the large language model to enhance the communication and cooperation efficiency between multi-agents, optimizes task allocation and collaborative decision-making through task requirement semantic mapping; introduces a concept drift detection and dynamic adjustment mechanism to improve the long-term adaptability in complex dynamic environments.

[0056] The multi-agent task collaboration method in the embodiment of the present invention has been described above. Next, the multi-agent task collaboration device in the embodiment of the present invention will be described. The multi-agent task collaboration device is referred to Figure 2 , an embodiment of the multi-agent task collaboration device in the embodiment of the present invention includes: A data acquisition module 201, configured to acquire visual images, text descriptions, voice instructions, and structured sensor signals, and perform spatio-temporal alignment and format conversion processing on the acquired data to obtain a multi-modal data set; The semantic encoding module 202 is used to perform deep fusion and unified semantic encoding on various modal inputs of the collected multimodal dataset by using a multimodal large language model, so as to obtain high-dimensional cross-modal feature embeddings and semantic representation files; The task coordination module 203 is used to correspond the local task requirements of multiple agents with cross-modal semantics based on the high-dimensional cross-modal feature embeddings and semantic representation files, obtain a task requirement semantic mapping, and perform task division and time arrangement of the multiple agents according to the task requirement semantic mapping, so as to obtain a multi-agent cooperation strategy file; The dynamic adjustment module 204 is used to, when the multiple agents execute the multi-agent cooperation strategy file, perform real-time sampling on key data and operation results during the execution process, compare the sampled data with the high-dimensional semantic representation to obtain a concept drift detection result, and when the concept drift detection result indicates that a concept has a progressive drift, perform dynamic adjustment on the multimodal large language model.

[0057] In the embodiment of the present invention, the multi-agent task cooperation device runs the above multi-agent task cooperation method. The multi-agent task cooperation device performs deep fusion and unified semantic encoding on the collected multimodal dataset by using a multimodal large language model to obtain high-dimensional cross-modal feature embeddings and semantic representation files, which are used to correspond the local task requirements of multiple agents with cross-modal semantics, obtain a task requirement semantic mapping, and perform task division and time arrangement; when the multiple agents execute tasks, perform real-time sampling on key data and operation results, compare them with the high-dimensional semantic representation to obtain a concept drift detection result, and when it is detected that a concept has a progressive drift, perform dynamic adjustment on the multimodal large language model. This method uses a large language model to enhance the communication and cooperation efficiency among multiple agents, optimizes task allocation and cooperation decision-making through task requirement semantic mapping; introduces a concept drift detection and dynamic adjustment mechanism to improve long-term adaptability in complex dynamic environments.

[0058] Above Figure 2 The multi-agent task cooperation device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Next, the multi-agent task cooperation device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0059] Figure 3FIG. 0 is a schematic structural diagram of a multi-agent task collaboration device provided by an embodiment of the present invention. The multi-agent task collaboration device 300 may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device terminals). Among them, the memory 320 and the storage medium 330 may be transient storage or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the multi-agent task collaboration device 300. Further, the processor 310 may be configured to communicate with the storage medium 330 and execute a series of instruction operations in the storage medium 330 on the multi-agent task collaboration device 300 to implement the steps of the above multi-agent task collaboration method.

[0060] The multi-agent task collaboration device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 The shown multi-agent task collaboration device structure does not limit the multi-agent task collaboration device provided by the present invention, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0061] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium, and may also be a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the multi-agent task collaboration method.

[0062] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system or device and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0063] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0064] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. A multi-agent task collaboration method, characterized in that: The multi-agent task collaboration method comprises: Collect visual images, text descriptions, voice commands, and structured sensor signals, and perform spatiotemporal alignment and format conversion on the collected data to obtain a multimodal dataset; Use a multimodal large language model to deeply fuse and unify the semantic encoding of various modal inputs of the collected multimodal dataset to obtain high-dimensional cross-modal feature embedding and semantic representation files; Based on the high-dimensional cross-modal feature embedding and semantic representation file, the local task requirements of the multi-agents are matched with the cross-modal semantics to obtain a task requirement semantic mapping, and according to the task requirement semantic mapping, the task division and time arrangement of the multi-agents are performed to obtain a multi-agent collaboration strategy file; When the multi-agent executes the multi-agent collaborative strategy file, the key data and operation results in the execution process are sampled in real time, and the sampled data is compared with the high-dimensional semantic representation to obtain a concept shift detection result. When the concept shift detection result is a gradual drift of the concept, the multimodal large language model is dynamically adjusted.

2. The multi-agent task collaboration method according to claim 1, characterized in that: The visual images, text descriptions, voice commands and structured sensor signals are collected, and the collected data are subjected to spatiotemporal alignment and format conversion processing to obtain a multimodal data set including: Performing convolutional neural network processing on the visual image to extract an image feature vector, applying a BERT model to embed words on the text description to obtain a text feature vector, extracting audio features from the voice command through MFCC transformation to obtain a voice feature vector, and performing wavelet transformation on the structured sensor signal to obtain a sensor feature vector; Aligning the image feature vector, the text feature vector, the speech feature vector, and the sensor feature vector according to timestamps, constructing a multimodal time series matrix, and applying a dynamic time warping algorithm to the multimodal time series matrix to obtain spatiotemporal alignment data; Mapping the data of different modes in the spatiotemporal alignment data to a unified numerical range, and reducing the feature dimension by principal component analysis to obtain a feature matrix after dimension reduction; A tensor decomposition algorithm is applied to the feature matrix after dimensionality reduction to extract shared and independent modal features, and the features are integrated into a tensor structure in a unified format to obtain the multimodal dataset.

3. The multi-agent task collaboration method according to claim 1, characterized in that: The multimodal large language model is used to perform deep fusion and unified semantic encoding on various modal inputs of the collected multimodal data set to obtain high-dimensional cross-modal feature embedding and semantic representation files, including: Applying a multi-head attention mechanism to the multimodal dataset, calculating the attention weights within each modality and between each modality, and obtaining attention weighted features; Inputting the attention weighted features into a multi-layer Transformer encoder to perform deep feature fusion to obtain a fused feature representation; Applying a contrastive learning loss function to the fused feature representation to optimize the semantic consistency of the features and obtain a semantically aligned feature vector; The semantically aligned feature vectors are projected into a high-dimensional space through a nonlinear mapping layer, a feature relationship graph is constructed, and a high-dimensional cross-modal feature embedding and semantic representation file is obtained.

4. The multi-agent task collaboration method according to claim 1, characterized in that: Based on the high-dimensional cross-modal feature embedding and semantic representation file, the local task requirements of the multi-agents are matched with the cross-modal semantics to obtain the task requirement semantic mapping, and according to the task requirement semantic mapping, the task division and time arrangement of the multi-agents are performed to obtain the multi-agent collaboration strategy file, including: Applying the t-SNE algorithm to the high-dimensional cross-modal feature embedding and semantic representation file for dimensionality reduction visualization, and using the DBSCAN density clustering algorithm to cluster the dimensionality-reduced data into task types to obtain task category labels; According to the task category label and the original task description, a task-semantic bidirectional mapping table is constructed to obtain a task requirement semantic mapping, and based on the task requirement semantic mapping, the front-back dependencies and resource requirements between tasks are analyzed to construct a task dependency graph; The task dependency graph is divided by using a graph partitioning algorithm to obtain a preliminary task allocation plan. A genetic algorithm is applied to the preliminary task allocation plan to optimize the task execution sequence and time allocation, and generate a multi-agent collaboration strategy file containing the operation instructions and communication protocols of each agent.

5. The multi-agent task collaboration method according to claim 4, characterized in that: The step of constructing a task-semantics bidirectional mapping table according to the task category label and the original task description to obtain a task requirement semantic mapping, and analyzing the front-back dependencies and resource requirements between tasks based on the task requirement semantic mapping to construct a task dependency graph includes: Calculating semantic similarity between the task category label and the original task description to obtain a task-semantic association matrix; According to the task-semantic association matrix, a task-semantic bidirectional mapping table is constructed using a bidirectional maximum matching algorithm to obtain a task requirement semantic mapping; Based on the task requirement semantic mapping, extract the timing information and resource requirement keywords in the task description to obtain a task attribute vector; Using a graph neural network to model the task attribute vector, construct an initial task relationship graph, and obtain a task node representation; A graph attention mechanism is applied to the task node representation to calculate the dependency strength between tasks and obtain a task dependency graph.

6. The multi-agent task collaboration method according to claim 1, characterized in that: Based on the high-dimensional cross-modal feature embedding and semantic representation file, the local task requirements of the multi-agents are matched with the cross-modal semantics to obtain the task requirement semantic mapping, and according to the task requirement semantic mapping, the task division and time arrangement of the multi-agents are performed to obtain the multi-agent collaboration strategy file, including: Use the distributed stream processing framework to perform real-time sampling and feature extraction on key data of the multi-agent execution process to obtain real-time feature sequences; The Wasserstein distance between the real-time feature sequence and the preset high-dimensional semantic representation is calculated, and an exponentially weighted moving average method is applied for smoothing to obtain a concept shift metric value; Applying the CUSUM algorithm to the concept drift metric value to perform cumulative sum analysis, setting a dynamic threshold, and obtaining a concept drift detection result; When the concept drift detection result exceeds a dynamic threshold, incremental learning and knowledge distillation techniques are used to update the parameters of the multimodal large language model.

7. The multi-agent task collaboration method according to claim 6, characterized in that: The applying the CUSUM algorithm to the concept drift metric value to perform cumulative sum analysis, setting a dynamic threshold, and obtaining a concept drift detection result includes: Segmenting the concept shift metric value with a sliding window to obtain time series data segments; Using an adaptive kernel density estimation method to perform probability distribution modeling on the time series data segment to obtain a current concept distribution; According to the current concept distribution, a CUSUM statistic is calculated, and compared with the historical CUSUM statistic to obtain a cumulative deviation value; The accumulated deviation value is smoothed by using an exponentially weighted moving average algorithm to obtain a smoothed CUSUM curve; Based on the historical data of the smoothed CUSUM curve, a dynamic threshold is calculated using a quantile regression method to obtain an adaptive threshold function; The smoothed CUSUM curve is compared with the adaptive threshold function to determine the optimal segmentation point and obtain a concept drift detection result.

8. A multi-agent task coordination device, characterized in that: The multi-agent task coordination device comprises: The data acquisition module is used to collect visual images, text descriptions, voice commands, and structured sensor signals, and perform spatiotemporal alignment and format conversion on the collected data to obtain a multimodal data set; The semantic encoding module is used to use the multimodal large language model to perform deep fusion and unified semantic encoding of various modal inputs of the collected multimodal data set to obtain high-dimensional cross-modal feature embedding and semantic representation files; A task collaboration module, which is used to match the local task requirements of the multi-agents with the cross-modal semantics based on the high-dimensional cross-modal feature embedding and semantic representation file, obtain a task requirement semantic mapping, and perform task division and time arrangement of the multi-agents according to the task requirement semantic mapping to obtain a multi-agent collaboration strategy file; The dynamic adjustment module is used to sample the key data and operation results in real time during the execution of the multi-agent collaborative strategy file, and compare the sampled data with the high-dimensional semantic representation to obtain the concept shift detection result. When the concept shift detection result is that the concept has a gradual drift, the multimodal large language model is dynamically adjusted.

9. A multi-agent task collaboration device, characterized in that: The multi-agent task coordination device comprises: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory so that the multi-agent task collaboration device executes the steps of the multi-agent task collaboration method as described in any one of claims 1-7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the steps of the multi-agent task collaboration method as described in any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Multi-model agent collaboration method and system

    CN120235428A

  • A multi-model intelligent agent collaboration method and system

    CN120235428B

  • Cloud task scheduling method and device based on intelligent agent, equipment and medium

    CN120256065A

  • Digital employee collaborative screening method and system

    CN120633673A

  • Digital employee collaboration screening method and system

    CN120633673B