Multi-modal large model information cooperative processing method, device, equipment and medium

By combining cross-modal privacy awareness, federated learning, and knowledge graphs, the shortcomings of large multimodal models in privacy-related awareness and security monitoring are addressed, achieving more efficient cross-modal data processing and decision-making accuracy.

CN120708014APending Publication Date: 2025-09-26YANTAI CAIKE INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510954464.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing large multimodal models have defects in cross-modal privacy association perception and security monitoring. They are unable to effectively capture the dynamic privacy associations between text descriptions and heterogeneous data such as medical images. In addition, the lack of causal reasoning for security monitoring makes it difficult to trace the causal chain of cross-modal data tampering.

Method used

Through cross-modal privacy-aware permission label binding processing, standardized data streams with dynamic permission labels are generated. Dynamic task allocation is performed using a federated learning-driven adaptive task scheduling engine. Causal reasoning monitoring is performed in conjunction with an explainable security agent. Semantic alignment processing is performed through knowledge graphs to generate fused cognitive results.

Benefits of technology

It improves the cross-modal privacy association perception capability, enhances the causal reasoning capability of security monitoring, optimizes the efficiency of multi-model collaboration, and improves the accuracy of fusion decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708014A_ABST
    Figure CN120708014A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal large models. By providing the multi-modal large model information co-processing method and device, the equipment and the medium, the method comprises the following steps: carrying out cross-modal privacy perception permission label binding processing on input multi-modal data, and generating a standardized data stream with a dynamic permission label; performing dynamic task allocation processing on the standardized data flow through a federated learning driven adaptive task scheduling engine, and generating a multi-modal intermediate result including a cross-modal interaction data flow; performing causal reasoning monitoring processing on the cross-modal interaction data flow through an interpretable security agent to generate a real-time security situation report; and performing semantic alignment processing based on a knowledge graph on the multi-modal intermediate result to generate a fusion cognition result so as to realize the technical effects of improving cross-modal privacy association perception capability, enhancing causal reasoning capability of security monitoring, optimizing multi-model cooperation efficiency and improving fusion decision accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal large models, and in particular to a method, apparatus, device and medium for collaborative processing of multimodal large model information. Background Art

[0002] With the rapid development of artificial intelligence (AI), large multimodal models are playing an increasingly important role in fields such as intelligent medical diagnosis, financial risk management, and autonomous driving. Collaborative processing of multimodal information is a core technology for achieving cross-modal semantic understanding and a key breakthrough in improving the cognitive capabilities of large models. Effective collaborative processing aims to integrate heterogeneous data sources such as text, images, and audio to generate secure and accurate fusion decision-making results.

[0003] However, the collaborative processing of related multimodal large model information has the following problems: (1) Lack of cross-modal privacy association perception: static permission strategies cannot capture the dynamic privacy associations between text descriptions and heterogeneous data such as medical images; (2) Lack of causal reasoning for security monitoring: rule matching mechanisms make it difficult to trace the causal chain of cross-modal data tampering. Summary of the Invention

[0004] Based on this, it is necessary to provide multimodal large model information collaborative processing methods, devices, equipment and media to address the above-mentioned technical problems, so as to achieve the technical effects of improving cross-modal privacy association perception capabilities, enhancing causal reasoning capabilities of security monitoring, optimizing multi-model collaborative efficiency and improving the accuracy of fusion decision-making.

[0005] In a first aspect, the present application provides a method for collaborative processing of multimodal large model information, the method comprising:

[0006] Perform cross-modal privacy-aware permission label binding on the input multimodal data to generate a standardized data stream with dynamic permission labels;

[0007] Based on collaborative task requirements, a federated learning-driven adaptive task scheduling engine dynamically allocates tasks to standardized data streams, generating multimodal intermediate results including cross-modal interactive data streams.

[0008] Through an explainable security agent, causal reasoning monitoring and processing are performed on cross-modal interactive data streams to generate real-time security situation reports. The security situation reports generated by causal reasoning monitoring and processing are fed back to the adaptive task scheduling engine in real time, forming a closed-loop security control mechanism.

[0009] Perform semantic alignment processing on the multimodal intermediate results based on the knowledge graph to generate fused cognitive results.

[0010] Furthermore, based on collaborative task requirements, a federated learning-driven adaptive task scheduling engine dynamically allocates tasks to standardized data streams, generating multimodal intermediate results including cross-modal interactive data streams, including:

[0011] Perform cross-modal dependency analysis on collaborative task requirements to generate a task topology graph including semantic dependencies.

[0012] Based on the federated learning framework, the capabilities of multiple modal processing models participating in the collaboration are quantitatively evaluated to generate capability description vectors that describe the computational characteristics of the models.

[0013] According to the task topology map and capability description vector, dynamic task-model mapping is performed to generate a task allocation instruction set;

[0014] The task assignment instruction set is transmitted to the corresponding modal processing model through the encrypted task distribution channel, driving the model execution to generate multimodal intermediate results.

[0015] Furthermore, based on the federated learning framework, the capabilities of multiple modal processing models participating in the collaboration are quantitatively evaluated to generate capability description vectors that describe the computational characteristics of the models, including:

[0016] Send heterogeneous computing capability probes to each modality processing model through the federated evaluation task distributor to generate model response datasets;

[0017] Perform cross-platform computational feature extraction on the model response dataset to generate a feature matrix containing delay sensitivity and memory usage patterns;

[0018] The federated feature aggregator performs dimension reduction on the feature matrix to generate standardized capability indicators with unified dimensions.

[0019] The standardized capability indicators are mapped to coordinate points in the multidimensional vector space to generate capability description vectors.

[0020] Furthermore, the task assignment instruction set is transmitted to the corresponding modal processing model through the encrypted task distribution channel, driving the model execution to generate multimodal intermediate results, including:

[0021] Performing hierarchical homomorphic encryption on the task allocation instruction set to generate an encrypted instruction package including execution logic;

[0022] Directly transmit the encrypted instruction packet to the target modal processing model through the federated routing protocol and generate an instruction reception confirmation signal;

[0023] Based on the instruction reception confirmation signal, the target model performs security instruction decapsulation processing to generate executable task instructions;

[0024] Drive the modal processing model to run executable task instructions in a trusted execution environment to generate initial intermediate results;

[0025] Perform zero-knowledge verification on the initial intermediate results to generate multimodal intermediate results.

[0026] Furthermore, an explainable security agent performs causal reasoning monitoring on cross-modal interaction data streams to generate real-time security situation reports, including:

[0027] Construct and process the interactive event graph of the cross-modal interaction data stream to generate a cross-modal interaction relationship graph;

[0028] Perform counterfactual causal reasoning based on cross-modal interaction relationship graphs to identify potential root cause modalities of abnormal interactions;

[0029] Perform multi-order impact propagation analysis on the identified root cause modes to generate a security threat chain path;

[0030] Perform dynamic risk assessment based on the security threat chain path to generate a real-time security situation report; the security situation report includes a threat path visualization map and a risk level assessment matrix.

[0031] Furthermore, counterfactual causal reasoning is performed based on the cross-modal interaction relationship graph to identify the potential root cause modalities of abnormal interactions, including:

[0032] Perform causal path separation on the cross-modal interaction relationship graph to generate a set of independent influence propagation paths;

[0033] Based on the set of independent impact propagation paths, virtual intervention operations are processed to generate counterfactual state simulation results;

[0034] Quantify the cross-modal influence of the counterfactual state simulation results to generate a modal influence weight vector;

[0035] Based on the modal influence weight vector, the root cause contribution is sorted and the potential root cause mode is generated.

[0036] Furthermore, the input multimodal data is subjected to cross-modal privacy-aware permission label binding processing to generate a standardized data stream with dynamic permission labels, including:

[0037] Perform cross-modal privacy correlation analysis on multimodal data to generate a privacy sensitivity heat map;

[0038] Based on the privacy sensitivity heat map, dynamic permission strategy generation is performed to generate transferable permission labels;

[0039] Based on transferable permission labels, the original multimodal data is desensitized in real time to generate a privacy-protected intermediate state of the data.

[0040] Perform cross-modal format collaborative conversion on the intermediate state of privacy-protected data to generate standardized data streams with dynamic permission labels.

[0041] In a second aspect, the present application further provides a multimodal large model information collaborative processing device, the device comprising:

[0042] The privacy-aware binding module is used to perform cross-modal privacy-aware permission label binding on the input multimodal data, generating a standardized data stream with dynamic permission labels;

[0043] The federated scheduling engine module is used to dynamically allocate tasks to standardized data streams based on collaborative task requirements through a federated learning-driven adaptive task scheduling engine, generating multimodal intermediate results including cross-modal interactive data streams;

[0044] The causal safety monitoring module is used to perform causal reasoning monitoring and processing on cross-modal interactive data streams through an explainable safety agent, generating real-time security situation reports. The security situation reports generated by the causal reasoning monitoring processing are fed back to the adaptive task scheduling engine in real time, forming a closed-loop security control mechanism.

[0045] The knowledge graph alignment module is used to perform semantic alignment processing on multimodal intermediate results based on the knowledge graph to generate fused cognitive results.

[0046] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any method in the first aspect of the present application are implemented.

[0047] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any method in the first aspect of the present application when the computer program is executed by a processor.

[0048] The present application provides a method, apparatus, equipment and medium for collaborative processing of multimodal large model information. The method includes: performing cross-modal privacy-aware permission label binding processing on the input multimodal data to generate a standardized data stream with dynamic permission labels; based on the collaborative task requirements, dynamically assigning tasks to the standardized data stream through a federated learning-driven adaptive task scheduling engine to generate a multimodal intermediate result including a cross-modal interactive data stream; performing causal reasoning monitoring processing on the cross-modal interactive data stream through an explainable security agent to generate a real-time security situation report; wherein the security situation report generated by the causal reasoning monitoring processing is fed back to the adaptive task scheduling engine in real time to form a closed-loop security control mechanism; performing semantic alignment processing on the multimodal intermediate results based on the knowledge graph to generate a fusion cognitive result, so as to achieve the technical effects of improving the cross-modal privacy association perception capability, enhancing the causal reasoning capability of security monitoring, optimizing the multi-model collaboration efficiency and improving the accuracy of fusion decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0050] Figure 1 Flowchart of a multimodal large model information collaborative processing method according to one embodiment of the present invention;

[0051] Figure 2 A flowchart for performing a capability quantification evaluation process on multiple modal processing models participating in collaboration based on a federated learning framework and generating a capability description vector that describes the computational characteristics of the models in one embodiment of the present invention.

[0052] Figure 3 This is a structural diagram of a multimodal large model information collaborative processing device in one embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the above-mentioned purposes, features and advantages of the present application more clearly understood, the specific implementation methods of the present application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to fully understand the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the connotation of the application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0054] like Figure 1As shown, the present application provides a multimodal large model information collaborative processing method, the method comprising:

[0055] S101: Perform cross-modal privacy-aware permission label binding processing on the input multimodal data to generate a standardized data stream with dynamic permission labels.

[0056] Specifically, cross-modal privacy association analysis is performed on input multimodal data such as text, images, and audio. By analyzing the semantic associations, contextual dependencies, and potential privacy association points between different modal data, the privacy-sensitive information contained in each modal data and the strength of its association in cross-modal interactions are identified, thereby generating a privacy sensitivity heat map that can intuitively reflect the distribution of privacy sensitivity. Then, based on the privacy risk distribution characteristics presented by this heat map, combined with preset privacy protection rules and dynamic adjustment strategies, portable permission labels are generated that can be adaptively adjusted as data flow scenarios, interaction objects, and usage purposes change. These labels include clear definitions of permissions such as data access scope, usage methods, and dissemination restrictions.

[0057] Then, based on the specific requirements of the portable permission labels, the original multimodal data is subjected to targeted real-time desensitization processing, such as replacing sensitive identifiers in text, blurring private areas in images, and perturbing feature information in audio. This generates a privacy-preserving data intermediate state that combines privacy protection with data usability. The privacy-preserving data intermediate state is then converted to a cross-modal format, unifying the data structure, encoding method, and transmission protocol. This allows data from different modalities to flow compatibly in a collaborative processing environment, generating a standardized data stream with dynamic permission labels.

[0058] S102: Based on collaborative task requirements, a federated learning-driven adaptive task scheduling engine is used to dynamically allocate tasks to standardized data streams, generating multimodal intermediate results including cross-modal interactive data streams.

[0059] Specifically, in response to collaborative task requirements, the dependencies between different modal data in task execution are analyzed, the semantic associations and processing sequences between text, image, audio and other data are clarified, and a task topology map containing the logical connections between subtasks is constructed. Based on the federated learning framework, the capabilities of multiple modal processing models participating in the collaboration are evaluated. By sending heterogeneous computing probes to collect the response characteristics of the models, computing features such as latency and memory usage are extracted and standardized to generate a capability description vector that reflects the model's processing capabilities. Based on the matching degree between the requirements of each subtask in the task topology map and the capability description vector, dynamic mapping of tasks and models is performed to generate a task allocation instruction set that includes task allocation rules and execution priorities.

[0060] The instruction set undergoes hierarchical encryption and is transmitted to the corresponding modal processing model via a federated routing protocol. Instructions are decrypted and executed within a trusted execution environment, generating initial intermediate results. Zero-knowledge verification ensures the integrity of the results. The initial intermediate results output by each model are then aggregated, extracting the correlation information generated by the interaction of different modal data, generating cross-modal interaction data streams, and ultimately integrating them to generate a multimodal intermediate result that includes these cross-modal interaction data streams.

[0061] S103: Perform causal reasoning monitoring and processing on the cross-modal interactive data stream through an explainable security agent to generate a real-time security situation report; the security situation report generated by the causal reasoning monitoring processing is fed back to the adaptive task scheduling engine in real time to form a closed-loop security control mechanism.

[0062] Specifically, the explainable safety agent parses cross-modal interaction data streams, extracting interaction events, temporal relationships, and correlation features between different modal data. It then constructs a cross-modal interaction relationship graph that intuitively presents the interaction logic of each modal data. Based on this graph, it simulates virtual interventions on the relevant modal data using counterfactual reasoning techniques, analyzes the changes in interaction states before and after the intervention, and quantifies the influence of each modality on abnormal interactions, thereby identifying the potential root cause modality of the anomaly.

[0063] Focusing on the root cause modality, the system tracks its propagation path in cross-modal interactions, analyzing its multi-order impact on other modal data and the overall collaborative process. This generates a security threat chain path that clearly describes the security threat's diffusion chain. Based on this, a dynamic risk assessment is conducted, combining the threat path's impact range, diffusion speed, and potential harm level. A real-time security situation report is generated, including a visual threat path map and a risk level assessment matrix. This security situation report is then fed back to the adaptive task scheduling engine in real time. The engine adjusts task allocation strategies based on the risk information in the report, achieving dynamic security control of the cross-modal collaborative process and forming a closed-loop security control mechanism.

[0064] S104: Perform semantic alignment processing on the multimodal intermediate results based on the knowledge graph to generate a fused cognitive result.

[0065] Specifically, for the different modal information such as text, images, and audio contained in the multimodal intermediate results, the core semantic features of each modal data, including key elements such as entities, attributes, and relationships, are extracted. These elements are then mapped to the preset knowledge graph framework to generate semantic fragments corresponding to each modality. Based on the ontological structure and domain rules of the knowledge graph, cross-modal association analysis is performed on the semantic fragments of different modalities, identifying and matching semantically equivalent or closely related entities and relationships, eliminating semantic ambiguity and conflicts caused by modal differences.

[0066] Entity linking technology is used to match the semantic elements of each modality with existing nodes in the knowledge graph, supplementing or improving the entity attributes and relationship networks in the graph, and integrating multimodal semantic information under a unified knowledge framework. Semantic consistency verification is performed on the integrated knowledge graph to ensure the logical rationality and accuracy of cross-modal information fusion, generating fused cognitive results that comprehensively reflect the deep connections and overall cognition of multimodal data.

[0067] An embodiment of the present application also provides a method for collaborative processing of multimodal large model information, including: performing cross-modal privacy-aware permission label binding processing on the input multimodal data to generate a standardized data stream with dynamic permission labels; based on the collaborative task requirements, dynamically assigning tasks to the standardized data stream through a federated learning-driven adaptive task scheduling engine to generate a multimodal intermediate result including a cross-modal interactive data stream; performing causal reasoning monitoring processing on the cross-modal interactive data stream through an explainable security agent to generate a real-time security situation report; wherein the security situation report generated by the causal reasoning monitoring processing is fed back to the adaptive task scheduling engine in real time to form a closed-loop security control mechanism; performing semantic alignment processing on the multimodal intermediate results based on the knowledge graph to generate a fusion cognitive result, so as to achieve the technical effects of improving the cross-modal privacy association perception capability, enhancing the causal reasoning capability of security monitoring, optimizing the multi-model collaboration efficiency and improving the accuracy of fusion decision-making.

[0068] Furthermore, based on collaborative task requirements, a federated learning-driven adaptive task scheduling engine dynamically allocates tasks to standardized data streams, generating multimodal intermediate results including cross-modal interactive data streams, including:

[0069] Perform cross-modal dependency analysis on collaborative task requirements to generate a task topology graph including semantic dependencies.

[0070] Based on the federated learning framework, the capabilities of multiple modal processing models participating in the collaboration are quantitatively evaluated to generate capability description vectors that describe the computational characteristics of the models.

[0071] According to the task topology map and capability description vector, dynamic task-model mapping is performed to generate a task allocation instruction set;

[0072] The task assignment instruction set is transmitted to the corresponding modal processing model through the encrypted task distribution channel, driving the model execution to generate multimodal intermediate results.

[0073] Specifically, in response to the needs of collaborative tasks, the tasks are broken down into subtasks such as text processing, image analysis, and audio recognition. The cross-modal dependencies of each subtask in data input, processing logic, and result output are analyzed, and the semantic association rules and processing timing constraints between different modal data are clarified. Then, a task topology map is constructed that can intuitively present the hierarchical structure and mutual dependencies of subtasks, which includes a description of the semantic dependency relationships of the interactions between modal data.

[0074] Based on the federated learning framework, evaluation indicators covering dimensions such as computing efficiency, modal adaptability, and resource usage are selected, and capability detection tasks containing diverse test data are sent to each modal processing model participating in the collaboration. The response characteristics and performance parameters of the model during the processing process are collected, and this information is subjected to standardized extraction and aggregation analysis to generate a capability description vector that can quantitatively reflect the computing characteristics of the model. The vector dimension corresponds to the preset evaluation indicator system.

[0075] Based on the modal type, processing complexity, accuracy requirements and other characteristics of each subtask in the task topology map, combined with the model adaptability presented by the capability description vector, dynamic matching of tasks and models is performed through preset matching rules and optimization algorithms to determine the optimal execution model and resource allocation plan for each subtask, and generate a task allocation instruction set containing subtask identification, target model address, execution parameters, priority and other information.

[0076] An encryption protocol is used to perform layered encryption on the task assignment instruction set, building a dedicated encrypted task distribution channel. Directed transmission is performed according to the target model address in the instruction set, ensuring the integrity and confidentiality of the instruction. After receiving the instruction, the model loads a standardized data stream based on the instruction content and executes the corresponding processing flow, generating intermediate processing results for each modality. These results generate cross-modal interactive data streams during the interaction process, and are ultimately integrated into a multimodal intermediate result that includes the cross-modal interactive data streams.

[0077] like Figure 2 As shown in the figure, based on the federated learning framework, the capabilities of multiple modal processing models participating in the collaboration are quantitatively evaluated to generate capability description vectors that describe the computational characteristics of the models, including:

[0078] S201: Send heterogeneous computing capability probes to each modality processing model through the federated evaluation task distributor to generate a model response dataset;

[0079] S202: Perform cross-platform computational feature extraction on the model response dataset to generate a feature matrix including delay sensitivity and memory usage patterns;

[0080] S203: Perform dimension reduction processing on the feature matrix through the federated feature aggregator to generate standardized capability indicators with unified dimensions;

[0081] S204: Map the standardized capability indicators to coordinate points in a multidimensional vector space to generate a capability description vector.

[0082] Specifically, a federated evaluation task distributor is used to generate and send heterogeneous computing capability probes including diverse computing scenarios and data loads for different types of modal processing models participating in the collaboration, such as text, images, and audio. The above probes contain processing tasks of different complexities and resource occupancy test instructions. By collecting the processing process records, resource call status, and output feedback of each model after receiving the probe, a model response dataset that can reflect the actual operating status of the model is generated through integration.

[0083] Systematically extract cross-platform computing features from the model response dataset, analyze the changing patterns of the model's response speed when processing tasks, and the dynamic allocation pattern of computing resources contained in the data, and focus on extracting the model's latency sensitivity to changes in task complexity, as well as the model's memory usage fluctuation characteristics under different loads and the memory usage pattern of the allocation strategy. These features are structured and organized according to preset dimensions to construct a feature matrix containing multi-dimensional computing characteristics.

[0084] The feature matrix is ​​dimensionally reduced through the federated feature aggregator, and feature selection and transformation techniques are used to eliminate redundant features while retaining core computing characteristic indicators. At the same time, the dimensions of features of different dimensions are unified to eliminate the impact caused by differences in evaluation indicator units. A set of standardized capability indicators with unified measurement standards that can be directly compared are generated. The above indicators more comprehensively reflect the core capabilities of the model in terms of computing efficiency, resource adaptability, etc.

[0085] The standardized capability indicators are mapped to a preset multi-dimensional vector space. Each dimension corresponds to a core capability indicator. The coordinate position of each indicator in the vector space is determined according to its value. The capability description vector that can fully describe the computational characteristics of the model is generated by combining the coordinate points. This vector can intuitively reflect the differences and advantages of different models in computing capabilities.

[0086] Furthermore, the task assignment instruction set is transmitted to the corresponding modal processing model through the encrypted task distribution channel, driving the model execution to generate multimodal intermediate results, including:

[0087] Performing hierarchical homomorphic encryption on the task allocation instruction set to generate an encrypted instruction package including execution logic;

[0088] Directly transmit the encrypted instruction packet to the target modal processing model through the federated routing protocol and generate an instruction reception confirmation signal;

[0089] Based on the instruction reception confirmation signal, the target model performs security instruction decapsulation processing to generate executable task instructions;

[0090] Drive the modal processing model to run executable task instructions in a trusted execution environment to generate initial intermediate results;

[0091] Perform zero-knowledge verification on the initial intermediate results to generate multimodal intermediate results.

[0092] Specifically, a layered homomorphic encryption mechanism is used to process the core information contained in the task allocation instruction set, such as task objectives, execution parameters, modal data indexes, etc.: the structural framework of the instruction set is encrypted at the first layer to ensure the security of the transmission format, and the execution logic details are encrypted at the second layer to protect the core content of the task. An encrypted instruction package is generated that retains the operability of the instructions and ensures the confidentiality of the information, to ensure that the instructions cannot be parsed even if they are intercepted during transmission.

[0093] A dedicated encrypted task distribution channel is constructed based on a federated routing protocol. This protocol incorporates node authentication, dynamic path optimization, and transmission encryption verification mechanisms. By identifying the unique identification information of the target modal processing model, encrypted instruction packets are routed to the corresponding model node. Upon receiving the instruction packet, the target model generates an instruction receipt confirmation signal containing the receipt time and the instruction packet integrity check result. This signal is then fed back to the sender through the original channel, completing the closed-loop confirmation of the transmission link.

[0094] On the target model side, based on the preset decryption key and security protocol, the encrypted instruction package is subjected to layered decapsulation processing: the signature information of the instruction package is verified to confirm the legitimacy of the source, and then the instruction structure and execution logic are decrypted layer by layer, eliminating redundant information that may be introduced during the transmission process, and restoring the executable task instructions containing specific processing steps and data call rules to ensure that the instructions meet the operating environment requirements of the model.

[0095] The modal processing model is driven to load executable task instructions into a trusted execution environment (TEE). This environment establishes an independent secure computing space through hardware isolation and software permission control mechanisms to prevent unauthorized access or tampering of data during instruction execution. Within this environment, the model invokes the corresponding standardized data stream, performs feature extraction, modal conversion, and other processing operations according to the instruction logic, and generates initial intermediate results containing the intermediate states of each modality and processing result identifiers.

[0096] Zero-knowledge verification technology is used to verify the validity of initial intermediate results: Without obtaining the specific content of the results, the verifier uses preset verification rules and feature hash comparison to confirm the integrity, consistency, and compliance with the task requirements. After verification, the initial intermediate results of different modalities and their interactive correlation information are integrated to generate a multimodal intermediate result that includes cross-modal interactive data streams.

[0097] Furthermore, an explainable security agent performs causal reasoning monitoring on cross-modal interaction data streams to generate real-time security situation reports, including:

[0098] Construct and process the interactive event graph of the cross-modal interaction data stream to generate a cross-modal interaction relationship graph;

[0099] Perform counterfactual causal reasoning based on cross-modal interaction relationship graphs to identify potential root cause modalities of abnormal interactions;

[0100] Perform multi-order impact propagation analysis on the identified root cause modes to generate a security threat chain path;

[0101] Perform dynamic risk assessment based on the security threat chain path to generate a real-time security situation report; the security situation report includes a threat path visualization map and a risk level assessment matrix.

[0102] Specifically, the explainable safety agent parses the cross-modal interaction data stream and extracts interaction events of different modal data such as text, images, and audio, including key information such as data transmission nodes, interaction timing, and content association types. The above events are abstracted into nodes and edges through the graph construction algorithm. The nodes represent the modal data or processing units involved in the interaction, and the edges represent the interaction relationships and characteristic attributes, thereby generating a cross-modal interaction relationship graph that can fully present the cross-modal data flow logic and association strength.

[0103] Based on the cross-modal interaction relationship graph, the counterfactual causal reasoning method is adopted. For the abnormal interaction nodes identified in the graph, virtual scenarios are simulated after removing specific modal nodes or changing their interaction characteristics. The changes in the interaction links between virtual scenarios and actual scenarios are compared, and the contribution of each modal node to the abnormal interaction is quantified. The modality with the highest contribution and at the starting position of the causal chain is selected as the potential root cause modality of the abnormal interaction.

[0104] Starting from the potential root cause modality, along the association path in the cross-modal interaction relationship map, trace the scope and degree of its influence on the downstream associated modalities: analyze the first-level associated modal directly affected by the root cause modal, and then expand layer by layer to the second-level and third-level associated modal affected by the first-level modal, record the abnormal performance and propagation path characteristics of the modalities at each level, and generate a security threat chain path that can clearly reflect the threat diffusion level and scope.

[0105] Based on the security threat chain path and combined with the preset risk assessment dimensions, the risk level of each node in the path is dynamically quantitatively assessed: corresponding risk weights are assigned to threat nodes at different levels, and an overall risk score is generated through weight aggregation. The threat path is presented in the form of a visual map, and an assessment matrix containing the risk level distribution in different scenarios is constructed. Finally, it is integrated into a real-time security situation report that includes a threat path visualization map and a risk level assessment matrix.

[0106] Furthermore, counterfactual causal reasoning is performed based on the cross-modal interaction relationship graph to identify the potential root cause modalities of abnormal interactions, including:

[0107] Perform causal path separation on the cross-modal interaction relationship graph to generate a set of independent influence propagation paths;

[0108] Based on the set of independent impact propagation paths, virtual intervention operations are processed to generate counterfactual state simulation results;

[0109] Quantify the cross-modal influence of the counterfactual state simulation results to generate a modal influence weight vector;

[0110] Based on the modal influence weight vector, the root cause contribution is sorted and the potential root cause mode is generated.

[0111] Specifically, for the cross-modal interaction relationship graph, a causal structure learning algorithm is used to perform causal path separation processing: identify modal node pairs with direct causal dependencies in the graph, construct a directed acyclic graph of causal relationships, and use a graph segmentation algorithm to decompose the graph into multiple independent sub-graphs. Each sub-graph represents a complete influence propagation path. The above path contains a complete causal chain from the initial triggering modality to the final result modality, and finally generates a set consisting of multiple independent causal paths.

[0112] Based on a set of independent influence propagation paths, virtual intervention operations are performed on the key modal nodes in each path: for specific modal nodes in the path, their situations in different states are simulated, and the impact of the above intervention on other nodes on the path and the final result is predicted through the counterfactual reasoning engine. The state changes and result deviations of each node in the path under each intervention scenario are recorded, and counterfactual state simulation results containing multiple virtual scenarios are generated.

[0113] The cross-modal impact of the counterfactual state simulation results is quantified: the degree and scope of the state changes of other modal nodes caused by the state change of a specific modal node in each virtual intervention scenario are analyzed, and characteristic indicators reflecting the intensity of the impact are extracted. The above indicators are mapped into an impact score with a unified dimension through standardization. The impact scores of each modality are combined to generate a modal influence weight vector that can reflect the intensity of the effect of different modalities in abnormal interactions.

[0114] Based on the modal influence weight vector, the root cause contribution of each modality is ranked from high to low according to the influence score: the modalities with influence scores significantly higher than the threshold are screened out, and a comprehensive evaluation is performed based on their position in the causal path. The modal that contributes most to the abnormal interaction is determined as the potential root cause modal, and a result set containing the root cause modal identification and its contribution evaluation is generated.

[0115] Furthermore, the input multimodal data is subjected to cross-modal privacy-aware permission label binding processing to generate a standardized data stream with dynamic permission labels, including:

[0116] Perform cross-modal privacy correlation analysis on multimodal data to generate a privacy sensitivity heat map;

[0117] Based on the privacy sensitivity heat map, dynamic permission strategy generation is performed to generate transferable permission labels;

[0118] Based on transferable permission labels, the original multimodal data is desensitized in real time to generate a privacy-protected intermediate state of the data.

[0119] Perform cross-modal format collaborative conversion on the intermediate state of privacy-protected data to generate standardized data streams with dynamic permission labels.

[0120] Specifically, cross-modal privacy association analysis is performed on input multimodal data such as text, images, and audio: the semantic associations and contextual dependencies between different modal data are analyzed, the privacy-sensitive elements contained in each modality are identified, and the correlation strength and leakage risk of the above elements in cross-modal interactions are evaluated. The distribution characteristics of privacy sensitivity are converted into a privacy sensitivity heat map through heat map visualization technology.

[0121] Based on the risk distribution presented by the privacy sensitivity heat map, combined with preset privacy protection rules and dynamic adjustment logic, migratable permission labels are generated: the labels contain limitations on the scope of data access, constraints on usage methods, and identification of the strength of privacy protection, and can be adaptively updated as the data flow scenario changes.

[0122] According to the specific requirements of the migratable permission labels, the original multimodal data is desensitized in real time: replacement or masking is used for sensitive fields in the text, blurring or pixel perturbation is performed on the private areas in the image, and spectral adjustment is performed on the feature information in the audio. At the same time, the integrity and availability of non-sensitive information are retained, generating a privacy-preserving data intermediate state that is both privacy-preserving and data-valid.

[0123] Collaborative cross-modal format conversion of privacy-protected data intermediate states: Unify the structural specifications and transmission protocols of data in different modalities to ensure that text, images, audio and other data can flow compatibly in a collaborative processing environment. At the same time, dynamic permission tags are embedded in the converted unified format to generate standardized data streams with dynamic permission tags.

[0124] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0125] In one embodiment, if Figure 3 As shown, the present application also provides a multimodal large model information collaborative processing device 300, which includes:

[0126] The privacy-aware binding module 301 is used to perform cross-modal privacy-aware permission tag binding processing on the input multimodal data to generate a standardized data stream with dynamic permission tags;

[0127] A federated scheduling engine module 302 is configured to dynamically allocate tasks to standardized data streams based on collaborative task requirements using a federated learning-driven adaptive task scheduling engine, generating multimodal intermediate results including cross-modal interactive data streams.

[0128] The causal safety monitoring module 303 is used to perform causal reasoning monitoring on cross-modal interactive data streams through an explainable safety agent to generate real-time security situation reports. The security situation reports generated by the causal reasoning monitoring process are fed back to the adaptive task scheduling engine in real time, forming a closed-loop security control mechanism.

[0129] The knowledge graph alignment module 304 is used to perform semantic alignment processing on the multimodal intermediate results based on the knowledge graph to generate a fused cognitive result.

[0130] Specifically, the privacy-aware binding module 301 receives multimodal input data such as text, images, and audio, analyzes the privacy associations between different modalities, identifies privacy-sensitive elements in each modality and evaluates their association strength, and generates a privacy sensitivity heat map showing the distribution of sensitivity. Based on the heat map and scenario-based privacy rules, a portable permission label containing access scope, usage constraints, and protection strength is generated. Targeted desensitization is performed on the original data based on the label to generate a data intermediate state with privacy protection. The structural specifications and transmission protocols of cross-modal data are unified, and dynamic permission labels are embedded therein to generate a standardized data stream with dynamic permission labels.

[0131] The federated scheduling engine module 302 accepts standardized data streams, analyzes collaborative task requirements to clarify cross-modal dependencies between subtasks, and constructs a task topology map containing semantic associations and processing sequences. Based on the federated learning framework, it then sends capability detection tasks to each modal processing model participating in the collaboration, collects response features, extracts computational characteristics, and generates a quantized capability description vector. Subsequently, it dynamically matches the subtask features of the task topology map with the capability description vector to generate a task allocation instruction set containing the target model, execution parameters, and priority. The instructions are transmitted to the corresponding model through an encrypted channel, driving the model to execute processing in a trusted environment and generate multimodal intermediate results including cross-modal interaction data streams.

[0132] The causal security monitoring module 303 parses cross-modal interaction data streams, extracts interaction events, temporal relationships, and correlation features, and constructs a cross-modal interaction relationship graph. Using counterfactual reasoning techniques based on this graph, it simulates virtual interventions to identify potential root-cause modalities of abnormal interactions. It then tracks the impact propagation paths of the root-cause modalities, analyzes their effects on multi-level correlated modalities, and generates a chain path for security threats. Dynamic risk assessments are conducted based on factors such as threat diffusion speed and impact range, generating a real-time security situation report containing a visual threat path graph and a risk level assessment matrix. This report is fed back to the federated scheduling engine module 302 in real time, forming a closed-loop security control mechanism to dynamically adjust task allocation.

[0133] The knowledge graph alignment module 304 receives the multimodal intermediate results, extracts the core semantic features such as entities, attributes, and relationships of each modality, and maps them to the preset knowledge graph framework to form semantic fragments. Based on the knowledge graph ontology structure and domain rules, cross-modal association analysis is performed on the semantic fragments of different modalities, matching equivalent or closely related entities and relationships to eliminate semantic ambiguity. Through entity linking technology, the semantic elements of each modality are matched with the knowledge graph nodes, and the entity attributes and relationship network are improved. After semantic consistency verification, a fusion cognitive result is generated that comprehensively reflects the deep associations of multimodal data.

[0134] The federated scheduling engine module 302 is further configured to:

[0135] Perform cross-modal dependency analysis on collaborative task requirements to generate a task topology graph including semantic dependencies.

[0136] Based on the federated learning framework, the capabilities of multiple modal processing models participating in the collaboration are quantitatively evaluated to generate capability description vectors that describe the computational characteristics of the models.

[0137] According to the task topology map and capability description vector, dynamic task-model mapping is performed to generate a task allocation instruction set;

[0138] The task assignment instruction set is transmitted to the corresponding modal processing model through the encrypted task distribution channel, driving the model execution to generate multimodal intermediate results.

[0139] The federated scheduling engine module 302 is further configured to:

[0140] Send heterogeneous computing capability probes to each modality processing model through the federated evaluation task distributor to generate model response datasets;

[0141] Perform cross-platform computational feature extraction on the model response dataset to generate a feature matrix containing delay sensitivity and memory usage patterns;

[0142] The federated feature aggregator performs dimension reduction on the feature matrix to generate standardized capability indicators with unified dimensions.

[0143] The standardized capability indicators are mapped to coordinate points in the multidimensional vector space to generate capability description vectors.

[0144] The federated scheduling engine module 302 is further configured to:

[0145] Performing hierarchical homomorphic encryption on the task allocation instruction set to generate an encrypted instruction package including execution logic;

[0146] Directly transmit the encrypted instruction packet to the target modal processing model through the federated routing protocol and generate an instruction reception confirmation signal;

[0147] Based on the instruction reception confirmation signal, the target model performs security instruction decapsulation processing to generate executable task instructions;

[0148] Drive the modal processing model to run executable task instructions in a trusted execution environment to generate initial intermediate results;

[0149] Perform zero-knowledge verification on the initial intermediate results to generate multimodal intermediate results.

[0150] The causal safety monitoring module 303 is further configured to:

[0151] Construct and process the interactive event graph of the cross-modal interaction data stream to generate a cross-modal interaction relationship graph;

[0152] Perform counterfactual causal reasoning based on cross-modal interaction relationship graphs to identify potential root cause modalities of abnormal interactions;

[0153] Perform multi-order impact propagation analysis on the identified root cause modes to generate a security threat chain path;

[0154] Perform dynamic risk assessment based on the security threat chain path to generate a real-time security situation report; the security situation report includes a threat path visualization map and a risk level assessment matrix.

[0155] The causal safety monitoring module 303 is further configured to:

[0156] Perform causal path separation on the cross-modal interaction relationship graph to generate a set of independent influence propagation paths;

[0157] Based on the set of independent impact propagation paths, virtual intervention operations are processed to generate counterfactual state simulation results;

[0158] Quantify the cross-modal influence of the counterfactual state simulation results to generate a modal influence weight vector;

[0159] Based on the modal influence weight vector, the root cause contribution is sorted and the potential root cause mode is generated.

[0160] The privacy-aware binding module 301 is further configured to:

[0161] Perform cross-modal privacy correlation analysis on multimodal data to generate a privacy sensitivity heat map;

[0162] Based on the privacy sensitivity heat map, dynamic permission strategy generation is performed to generate transferable permission labels;

[0163] Based on transferable permission labels, the original multimodal data is desensitized in real time to generate a privacy-protected intermediate state of the data.

[0164] Perform cross-modal format collaborative conversion on the intermediate state of privacy-protected data to generate standardized data streams with dynamic permission labels.

[0165] In one embodiment, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0166] In one embodiment, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above-mentioned method embodiments when the computer program is executed by a processor.

[0167] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0168] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.

Claims

1. A multimodal large model information collaborative processing method, characterized in that: The method comprises: Perform cross-modal privacy-aware permission label binding on the input multimodal data to generate a standardized data stream with dynamic permission labels; Based on collaborative task requirements, a federated learning-driven adaptive task scheduling engine is used to dynamically allocate tasks to the standardized data stream, generating multimodal intermediate results including cross-modal interactive data streams; Performing causal reasoning monitoring on the cross-modal interactive data stream through an explainable security agent to generate a real-time security situation report; wherein the security situation report generated by the causal reasoning monitoring process is fed back to the adaptive task scheduling engine in real time to form a closed-loop security control mechanism; The multimodal intermediate results are subjected to semantic alignment processing based on the knowledge graph to generate a fused cognitive result.

2. The multimodal large model information collaborative processing method according to claim 1, characterized in that: Based on the collaborative task requirements, the adaptive task scheduling engine driven by federated learning performs dynamic task allocation processing on the standardized data stream to generate a multimodal intermediate result including a cross-modal interactive data stream, including: Perform cross-modal dependency analysis on collaborative task requirements to generate a task topology graph including semantic dependencies. Based on the federated learning framework, the capabilities of multiple modal processing models participating in the collaboration are quantitatively evaluated to generate capability description vectors that describe the computational characteristics of the models. Performing task-model dynamic mapping processing according to the task topology map and the capability description vector to generate a task allocation instruction set; The task allocation instruction set is transmitted to the corresponding modal processing model through an encrypted task distribution channel, driving the model execution to generate the multimodal intermediate result.

3. The multimodal large model information collaborative processing method according to claim 2, characterized in that: Based on the federated learning framework, the capability quantification evaluation process is performed on multiple modal processing models participating in the collaboration to generate capability description vectors that describe the computational characteristics of the models, including: Send heterogeneous computing capability probes to each modality processing model through the federated evaluation task distributor to generate model response datasets; Performing cross-platform computational feature extraction processing on the model response dataset to generate a feature matrix including delay sensitivity and memory usage patterns; Performing dimension reduction processing on the feature matrix through a federated feature aggregator to generate standardized capability indicators with unified dimensions; The standardized capability index is mapped to coordinate points in a multidimensional vector space to generate the capability description vector.

4. The multimodal large model information collaborative processing method according to claim 2, characterized in that: The step of transmitting the task allocation instruction set to the corresponding modal processing model through the encrypted task distribution channel and driving the model to execute and generate the multimodal intermediate result includes: Performing layered homomorphic encryption on the task assignment instruction set to generate an encrypted instruction packet including execution logic; Directly transmitting the encrypted instruction packet to the target modal processing model through a federated routing protocol, and generating an instruction reception confirmation signal; Based on the instruction reception confirmation signal, performing security instruction decapsulation processing on the target model side to generate executable task instructions; Driving the modal processing model to execute the executable task instructions in the trusted execution environment to generate an initial intermediate result; Perform zero-knowledge verification processing on the initial intermediate result to generate the multimodal intermediate result.

5. The multimodal large model information collaborative processing method according to claim 1, characterized in that: The process of performing causal reasoning monitoring on the cross-modal interaction data stream by using an explainable security agent to generate a real-time security situation report includes: Construct and process the interactive event graph of the cross-modal interaction data stream to generate a cross-modal interaction relationship graph; Performing counterfactual causal reasoning based on the cross-modal interaction relationship graph to identify potential root cause modalities of abnormal interactions; Performing multi-order impact propagation analysis on the identified root cause mode to generate a security threat chain path; Dynamic risk assessment processing is performed based on the security threat chain path to generate the real-time security situation report; wherein the security situation report includes a threat path visualization map and a risk level assessment matrix.

6. The multimodal large model information collaborative processing method according to claim 5, characterized in that: The counterfactual causal reasoning process is performed based on the cross-modal interaction relationship graph to identify the potential root cause modality of the abnormal interaction, including: Performing causal path separation processing on the cross-modal interaction relationship graph to generate a set of independent influence propagation paths; Based on the set of independent impact propagation paths, a virtual intervention operation is performed to generate a counterfactual state simulation result; Performing cross-modal influence quantification on the counterfactual state simulation result to generate a modal influence weight vector; Based on the modal influence weight vector, a root cause contribution ranking process is performed to generate the potential root cause modality.

7. The multimodal large model information collaborative processing method according to claim 1, characterized in that: The cross-modal privacy-aware permission tag binding process is performed on the input multimodal data to generate a standardized data stream with dynamic permission tags, including: Performing cross-modal privacy association analysis on the multimodal data to generate a privacy sensitivity heat map; Based on the privacy sensitivity heat map, dynamic permission strategy generation processing is performed to generate a migratable permission label; Based on the transferable permission labels, the original multimodal data is desensitized in real time to generate an intermediate state of data with privacy protection; The privacy-protected data intermediate state is subjected to cross-modal format collaborative conversion processing to generate the standardized data stream with dynamic permission labels.

8. A multimodal large model information collaborative processing device, characterized in that: The device comprises: The privacy-aware binding module is used to perform cross-modal privacy-aware permission label binding on the input multimodal data, generating a standardized data stream with dynamic permission labels; A federated scheduling engine module is used to dynamically allocate tasks to the standardized data stream based on collaborative task requirements through a federated learning-driven adaptive task scheduling engine, generating multimodal intermediate results including cross-modal interactive data streams; A causal safety monitoring module is configured to perform causal reasoning monitoring on the cross-modal interactive data stream through an explainable safety agent to generate a real-time security situation report; wherein the security situation report generated by the causal reasoning monitoring process is fed back to the adaptive task scheduling engine in real time to form a closed-loop security control mechanism; The knowledge graph alignment module is used to perform semantic alignment processing on the multimodal intermediate results based on the knowledge graph to generate a fused cognitive result.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multimodal large model information collaborative processing method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal large model information collaborative processing method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Domestic PC terminal cross-modal natural interaction method and system fusing generative AI

    CN121029005A

  • Artificial intelligence-driven cognitive disorder early screening and intervention integrated platform

    CN121583567A

  • Artificial intelligence-driven integrated platform for early screening and intervention of cognitive impairment

    CN121583567B

  • Dynamic data security protection method and device based on adversarial training under multi-modal large model

    CN121690761A