Multi-modal data management system and method and related product

By implementing data standardization, model scheduling, isolation classification, and risk compliance review modules in the multimodal data management system, the consistency and compliance issues in multimodal data processing have been resolved, enabling efficient and stable multimodal data processing that is adaptable to high-concurrency, low-latency production scenarios.

CN121919908APending Publication Date: 2026-04-24BEIJING XUEDIRUANJIAN DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XUEDIRUANJIAN DEVELOPMENT CO LTD
Filing Date
2026-01-06
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, the access and processing links for multimodal data are scattered, lacking a unified orchestration and feature alignment mechanism. This results in poor consistency in cross-modal data processing, significant latency jitter, and difficulty in meeting the stable experience requirements under high-concurrency scenarios. At the same time, the reasoning process and the final answer share the same transmission channel, making it difficult to accurately locate risk points during the review process. This leads to a high rate of false blocking and false rejection of data transmission, making it difficult to guarantee compliance.

Method used

The multimodal data management system, including a data standardization construction module, a model scheduling module, and an isolation, classification, and risk compliance audit module, enables unified encapsulation and standardized format verification of multimodal data. It determines the appropriate data processing model based on data characteristics and performs load balancing processing, isolates and classifies data, and combines it with a preset risk feature library for risk screening and compliance verification.

Benefits of technology

It improves the consistency, efficiency, and compliance of multimodal data processing, adapts to the needs of high-concurrency, low-latency production scenarios, optimizes data processing efficiency and system concurrency capacity, reduces the false and missed rates of risk identification, and ensures data compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121919908A_ABST
    Figure CN121919908A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data management system and method and related products, and the system comprises a data standardization construction module which is used for carrying out the unified packaging processing of multi-modal data and format standardization verification of multi-modal initial data, so as to generate standardized multi-modal data; the model scheduling module is used for determining a data processing model matched with the standardized multi-modal data based on the data characteristics of the standardized multi-modal data and executing load balancing processing on the data processing model matched with the same data characteristics so as to generate multi-source collaborative output data; and the isolation classification and risk compliance auditing module performs isolation classification processing on the multi-source collaborative output data and performs risk screening and compliance verification on the multi-source collaborative output data after the isolation classification processing in combination with a preset risk feature library so as to output outgoing data meeting compliance requirements. According to the method and the device, the consistency, the efficiency and the compliance of multi-modal data processing are improved, and the production scene requirements of high concurrency and low time delay are effectively met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a multimodal data management system, method, and related products. Background Technology

[0002] With the rapid development of digital technology, the application scenarios of multimodal data such as text and images are becoming increasingly widespread. Especially in high-concurrency, low-latency production scenarios, higher requirements are placed on the processing efficiency, compliance, and stability of multimodal data. For example, in scenarios such as intelligent learning devices and online service platforms, it is necessary to respond quickly to multimodal input requests while ensuring the compliance, security, and reliability of output data. This requires an efficient multimodal data management mechanism to support business operations.

[0003] With the development of multimodal technology, related data processing solutions have gradually emerged. In existing technical solutions, independent components such as text processing modules and image processing modules are deployed separately to deal with multimodal input. Different modal data are transmitted to the corresponding processing modules through their respective access links, and the results are directly integrated and output after processing. At the same time, a fixed keyword matching method is used for risk control, and some operational data is collected through distributed indicator collection tools for monitoring.

[0004] However, this existing technical solution has significant technical shortcomings in practical applications. The access and processing links for multimodal data are fragmented, lacking a unified orchestration and feature alignment mechanism, resulting in poor consistency in cross-modal data processing, significant latency jitter, and difficulty in meeting the stable experience requirements of high-concurrency scenarios. Furthermore, the reasoning process and the final answer share the same transmission channel, making it difficult to accurately locate risk points during the review process. This leads to a high rate of false and false interceptions in data transmission, making compliance difficult to guarantee. This technical shortcoming has become a key bottleneck restricting the reliable application of multimodal data in demanding production scenarios. Summary of the Invention

[0005] To address the aforementioned technical problems, this application provides a multimodal data management system, method, and related products to at least alleviate these problems.

[0006] The technical solutions provided in this application are as follows: A multimodal data management system, comprising: The data standardization construction module is used to perform unified encapsulation and format standardization verification of the initial multimodal data to generate standardized multimodal data. The model scheduling module determines a data processing model that is compatible with the standardized multimodal data based on the data characteristics of the standardized multimodal data, and performs load balancing processing on the data processing models that are compatible with the same data characteristics to generate multi-source collaborative output data. The isolation classification and risk compliance audit module isolates and classifies the multi-source collaborative output data, and performs risk screening and compliance verification on the isolated and classified multi-source collaborative output data in combination with a preset risk feature library, so as to output outgoing data that meets compliance requirements.

[0007] Optionally, the data standardization construction module includes: a multimodal data unified encapsulation unit, which encapsulates the initial multimodal data based on the field composition rules of data type, multimodal data standardization encoding, and business category identifier to form structured multimodal data; and a format standardization verification unit, which performs format standardization verification on the structured multimodal data and outputs standardized multimodal data that conforms to the format standardization verification.

[0008] Optionally, the model scheduling module includes: a data feature acquisition and model mapping unit, which determines a target data processing model set based on a preset mapping rule between the data processing model and the data type and business category of the standardized multimodal data; a model status evaluation unit, which generates status evaluation data for each data processing model in the target data processing model set based on the real-time running status data of each data processing model in the target data processing model set; and a load balancing control and data distribution unit, which performs load balancing processing on data processing models that adapt to the same data features based on the status evaluation data of each data processing model, and distributes the standardized multimodal data to the target data processing model set to generate multi-source collaborative output data.

[0009] Optionally, the isolation classification and risk compliance audit module includes: a dual-channel isolation processing unit, which performs dual-channel isolation processing on the multi-source collaborative output data based on a dual-channel transmission protocol of inference data channel and result data channel to obtain the inference data and the result data; and an independent identifier configuration unit, which configures inference identifier field and result identifier field for the inference data and the result data to generate inference data and result data with independent identifiers.

[0010] Optionally, the data standardization construction module, the model scheduling module, and the isolation audit and compliance processing module have an integrated interface layer structure, which constructs a communication mode with unified access specifications, real-time interaction, and streaming output.

[0011] Optionally, it also includes a model fault tolerance module, specifically including at least one of the following: a configuration strategy management unit, which configures the running configuration parameters of the data processing model differently according to different business categories, the running configuration parameters including identification information, priority, timeout threshold and maximum number of retries; an anomaly detection and automatic recovery unit, which performs real-time status detection on the data processing model matching the running configuration parameters based on the running configuration parameters configured differently by the configuration strategy management unit according to different business categories, synchronously collects running log statistics on the success rate and response latency data of the model and performs dynamic optimization processing, and if an anomaly is detected in the data processing model, performs self-recovery processing on the data processing model with an anomaly through a service anomaly status reset mechanism, and synchronously feeds back the self-recovery result and anomaly status data to the model switching unit; and a model switching unit, which receives the self-recovery result and anomaly status data fed back by the anomaly detection and automatic recovery unit, and, in combination with the running configuration parameters configured differently by the configuration strategy management unit according to different business categories, triggers a backup data processing model or backup multimodal data processing node adapted to the currently anomaly data processing model.

[0012] Optionally, it also includes a full-link monitoring module, which comprises at least one of a distributed tracing unit, a performance data acquisition unit, a multi-dimensional indicator statistics unit, and a data visualization unit. The distributed tracing unit embeds a unique tracking identifier into the entire process of the standardized multimodal data processing based on the full-link tracing mechanism to construct a full-process tracing link covering the data standardization construction module, the standardization construction module, and the isolation classification and risk compliance review module. The performance data acquisition unit collects request processing runtime data during the data request processing process based on the unique tracking identifier of the distributed tracing unit. The request processing runtime data includes response time and call success rate. The system receives the request processing operation data and error type data, and synchronously transmits the request processing operation data, carrying a unique identifier, to the multi-dimensional indicator statistics unit. The multi-dimensional indicator statistics unit, based on the unique tracking identifier of the distributed tracking unit, combines the data processing model dimension and business category dimension to calculate multi-dimensional business operation statistics data, including call volume, average response latency, percentage of degradation trigger times, and percentage of audit rule hit times. The data visualization unit receives the request processing operation data and the multi-dimensional business operation statistics data, and displays the request processing operation data and the multi-dimensional business operation statistics data in a structured dashboard format.

[0013] Optionally, it also includes a multi-layer risk control module, which specifically has at least one of a dual-review unit, a dynamic downgrade handling unit, and a streaming gating review unit: The dual-review unit includes an input-end risk identification unit and an output-end compliance verification unit. The input-end risk identification unit identifies pre-risk information such as risk identification and risk level for text-sensitive information and image-sensitive risks in the initial multimodal data. The output-end compliance verification unit performs secondary verification on the compliance data to be sent out. The dynamic downgrade handling unit collects the pre-risk information identification results, including risk identification and risk level, and combines the risk attributes of the data processing model to differentiate handling strategies based on risk level, so as to transmit the risk identification and handling strategies to the business layer and the streaming gating review unit. The streaming gating review unit receives the risk identification and handling strategies transmitted by the dynamic downgrade handling unit, performs streaming processing and real-time risk matching on the standardized multimodal data using a segmented review mode based on the handling strategies, and performs differentiated processing according to the matching results.

[0014] A multimodal data management method includes the following steps: performing unified multimodal data encapsulation and format standardization verification on initial multimodal data to generate standardized multimodal data; determining a data processing model adapted to the standardized multimodal data based on the data type and business category of the standardized multimodal data, and performing load balancing processing on the data processing models adapted to the same data type to generate multi-source collaborative output data; splitting the multi-source collaborative output data based on a dual-channel isolation mechanism to obtain inference data and result data, and assigning independent identifiers to them for differentiation; and performing risk screening on the inference data and compliance verification on the result data based on the independent identifiers and a preset risk feature library to output outgoing data that has been risk-filtered and meets compliance requirements.

[0015] An electronic device includes: a memory for storing executable program code; and a processor for calling and running the executable program code from the memory, thereby enabling the electronic device to implement the aforementioned multimodal data management method.

[0016] A computer storage medium storing at least one instruction or at least one program, wherein the at least one instruction or the at least one program is loaded and executed by a processor to implement the above-described multimodal data management method.

[0017] A computer program product or computer program, the computer program product or computer program including computer instructions, which, when executed by a processor, implement the above-described multimodal data management method.

[0018] The technical solution provided in this application has the following technical advantages: This application's solution achieves unified encapsulation and format verification of multimodal initial data through a data standardization construction module, eliminating format differences and compatibility issues between different modal data and laying a unified data foundation for cross-modal collaborative processing. The model scheduling module accurately matches and adapts models based on data features and performs load balancing, optimizing model resource configuration, avoiding single-node overload, and improving data processing efficiency and system concurrency capacity. The isolation classification and risk compliance audit module separates inference data from result data and conducts targeted risk screening and compliance verification, reducing audit interference and improving the targeting of risk identification and the reliability of compliance verification. Overall, it achieves improved consistency, efficiency, and compliance in multimodal data processing, effectively adapting to the needs of high-concurrency, low-latency production scenarios. Attached Figure Description

[0019] Figure 1 A block diagram of a multimodal data management system provided in one embodiment of this application; Figure 2 A flowchart of a multimodal data management method provided in one embodiment of this application; Figure 3 This is a structural block diagram of an electronic device provided in one embodiment of this application. Detailed Implementation

[0020] like Figure 1 As shown in the figure, this application provides a multimodal data management system, including: A multimodal data management system, comprising: The data standardization construction module is used to perform unified encapsulation and format standardization verification of the initial multimodal data to generate standardized multimodal data. The model scheduling module determines a data processing model that is compatible with the standardized multimodal data based on the data characteristics of the standardized multimodal data, and performs load balancing processing on the data processing models that are compatible with the same data characteristics to generate multi-source collaborative output data. The isolation classification and risk compliance audit module isolates and classifies the multi-source collaborative output data, and performs risk screening and compliance verification on the isolated and classified multi-source collaborative output data in combination with a preset risk feature library, so as to output outgoing data that meets compliance requirements.

[0021] Optionally, the data standardization construction module includes: a multimodal data unified encapsulation unit, which encapsulates the initial multimodal data based on the field composition rules of data type, multimodal data standardization encoding, and business category identifier to form structured multimodal data; and a format standardization verification unit, which performs format standardization verification on the structured multimodal data and outputs standardized multimodal data that conforms to the format standardization verification.

[0022] In implementing the above steps, to address the issues of heterogeneous data formats and fragmented processing links in multimodal interaction scenarios, an integrated design of "modal feature binding - structured encapsulation - multi-level verification" is adopted to transform initial multimodal data into standardized data. Unlike traditional methods that only perform basic format conversion, this application's solution deeply integrates data types, encoding rules, and business scenarios. Through customized field structures and verification logic, it ensures that the data not only meets unified processing requirements but also adapts to the scenario-based needs of subsequent model scheduling and risk control audits, providing a reliable data foundation for high-concurrency, low-latency multimodal services. This module consists of a unified multimodal data encapsulation unit and a format standardization verification unit. The former is responsible for the structured organization of the data, while the latter is responsible for the compliance verification of the data. The two units form a coherent "encapsulation-verification" process, ensuring the uniformity, integrity, and usability of the output data.

[0023] Preferably, the module implementation of the multimodal data unified encapsulation unit focuses on building an adaptive structured data organization mechanism. Through field semantic association and scenario-based adaptation of encoding rules, it solves the problem of the difficulty in uniformly processing multimodal data. First, it receives initial multimodal data input from the user. This data covers various forms such as text and images. Based on the designed modality recognition logic, it extracts the morphological features of the data, determines whether the data is text-based, image-based, or a text + image composite, and generates a corresponding modality type identifier (MessageType). This modality type identifier is directly associated with subsequent encoding rules and business processing strategies, ensuring the targeted nature of data processing. Context-specific standardized encoding is performed for data of different modalities. For text data, NFKC character standardization encoding rules are used for conversion to eliminate character encoding differences caused by different input devices (such as handwriting input on learning machines and external keyboard input), ensuring the consistency of text content. For image data, Exchangeable image file format (EXIF) information is first removed from the image to prevent leakage of user privacy data. Then, the image is uniformly converted to sRGB color space, and the image size is adjusted according to a preset ratio (the longest side is controlled within a reasonable range, such as not exceeding a specific value). It is saved as JPEG or PNG format with a preset compression quality (such as a compression ratio within a reasonable range). Finally, the image binary data is converted to Base64 encoding to ensure the compatibility of image data transmission between the client and the server. For text+ For image-composite data (such as essays with illustrations, or math problems containing both text and geometric figures), the text and image parts are processed using the corresponding encoding methods described above to generate modal association identifiers. These identifiers record the semantic correspondence between text and images in a business scenario (e.g., the association between a paragraph of essay text and its accompanying illustration, or the association between the text of a question and its corresponding geometric figure), ensuring the synergy of modal information in subsequent processing. Then, based on the learning scenario corresponding to the initial multimodal data, a corresponding business category identifier (BusinessID) is matched. This business category identifier is based on the core application scenarios preset by the learning machine, including math problem solving, essay correction, English oral practice, and explanation of natural science knowledge points. Each business category identifier corresponds to a specific model processing strategy and risk control rules, ensuring the scenario adaptability of data processing. Finally, following the fixed field composition rule of "modal type identifier - standardized coded data - modal association identifier (composite data only) - business category identifier", the above-processed data is encapsulated into structured multimodal data. This structured multimodal data is organized using a customized unified message structure (MultiModalMessage), with fixed field order and format to ensure rapid parsing and processing of the data by the subsequent model scheduling module and risk control module.

[0024] Preferably, the core of the format standardization verification unit module implementation is to construct a multi-level, scenario-based verification mechanism. By combining field-level verification with overall structure verification, it ensures that the structured multimodal data meets the format requirements for subsequent processing. Specifically, the complete set of fields for the structured multimodal data is first obtained, and each field is verified one by one based on preset field-level verification rules. For the modality type identifier field, it is verified whether it is a member of the preset set of legal values ​​(text, image, text + image composite). If an illegal identifier exists, the field verification is directly determined to have failed. For standardized encoded data fields, for text encoded data, it is verified whether it conforms to the NFKC encoding standard, whether there are garbled characters or unresolvable characters, and whether the text length is within a preset range (e.g., the length of a single text after encoding does not exceed a specific value) to avoid processing efficiency degradation caused by excessively long text. For image Base64 encoded data, it is verified whether its encoding format conforms to the Base64 syntax rules, and the validity of the encoding is verified through decoding tests. Simultaneously, it is verified whether the length of the encoded data is within a preset threshold range (e.g., the length of a single image Base64). The encoded length does not exceed a specific value, balancing transmission efficiency and image quality. For modal association identifier fields (only for composite data), the format is verified to match the preset association rules and to accurately associate the corresponding text-encoded data and image-encoded data, avoiding modal association errors. For business category identifier fields, it is verified to be a member of the preset set of legal business categories and to be consistent with the business attributes of the data content (e.g., the business category identifier corresponding to math problem image data cannot be the essay correction scenario identifier), ensuring the adaptability of subsequent model scheduling and risk control strategies. After field-level verification, an overall structure verification is performed to verify whether the field arrangement order of the structured multimodal data conforms to the preset encapsulation rules, whether the delimiters between fields are correct, and whether the overall data conforms to the format requirements of the unified message structure (MultiModalMessage), avoiding subsequent module parsing failures due to structural errors. For any abnormal data detected during the validation process, a validation anomaly report is generated, including the name of the abnormal field, the anomaly type (format error, length exceeding limit, illegal value, structural disorder, etc.), the location of the anomaly, and rectification suggestions. This validation anomaly report is fed back to the learning machine's data input end, and a data resubmission mechanism is triggered, requiring the data input end to correct the data according to the validation rules and resubmit it. For structured multimodal data that passes all validation items, it is marked as standardized multimodal data that conforms to the format standardization validation. This standardized multimodal data will be directly transmitted to the model scheduling module for subsequent processing to ensure the data reliability of subsequent processes.

[0025] Optionally, the model scheduling module includes: a data feature acquisition and model mapping unit, which determines a target data processing model set based on a preset mapping rule between the data processing model and the data type and business category of the standardized multimodal data; a model status evaluation unit, which generates status evaluation data for each data processing model in the target data processing model set based on the real-time running status data of each data processing model in the target data processing model set; and a load balancing control and data distribution unit, which performs load balancing processing on data processing models that adapt to the same data features based on the status evaluation data of each data processing model, and distributes the standardized multimodal data to the target data processing model set to generate multi-source collaborative output data.

[0026] This application's model scheduling module addresses the issues of insufficient model adaptation accuracy and low resource utilization in high-concurrency multimodal processing scenarios. Through a three-layer architecture design of "feature mapping - dynamic state evaluation - scenario-based load balancing," it achieves efficient collaboration between standardized multimodal data and data processing models. Unlike traditional scheduling methods that select models based on a single dimension or employ fixed load balancing strategies, this module deeply correlates data type, business category, and real-time model state. Through customized mapping rules and evaluation logic, it ensures scenario-specific model adaptation while improving resource utilization efficiency and real-time data processing, supporting the low-latency requirements of multimodal interaction. This module consists of a data feature acquisition and model mapping unit, a model state evaluation unit, and a load balancing control and data distribution unit working in tandem to form a coherent process of "feature extraction - model selection - state evaluation - load balancing - data distribution," ensuring the efficient generation of multi-source collaborative output data.

[0027] Preferably, the data feature acquisition and model mapping unit, which constructs a precise association mechanism between data and models, solves the problem of poor routing adaptability of traditional models through two-dimensional feature extraction and customized mapping rules. It first receives standardized multimodal data output by the data standardization construction module, and extracts core data features from the standardized multimodal data. The core data features include data type features and business category features. The data type features are determined based on the modality type identifier of the standardized multimodal data, specifying whether the data is text, image, or a text + image composite. The business category features are determined based on the business category identifier (BusinessID) of the standardized multimodal data, corresponding to specific application scenarios (such as solving math problems, grading essays, practicing English speaking, explaining natural science knowledge points, etc.). Subsequently, a pre-defined model mapping rule library is invoked. This rule library is built based on the actual needs of multimodal processing scenarios and adopts a two-dimensional relational table structure. The row dimension of the two-dimensional relational table is the data type (text, image, text + image composite), and the column dimension is the business category. The intersection of the row and column records information on all data processing models that adapt to the combination of data type and business category (including model identifier, model expertise, adaptation priority, etc.). Based on the extracted core data features, precise matching is performed in the two-dimensional relational table of the model mapping rule library to select data processing models that simultaneously adapt to the data type and business category features. For example, for pure text data, models focusing on logical deduction in the text model pool are matched; for text + image composite data, models supporting text content analysis and image correlation judgment in the multimodal model pool are matched. These selected data processing models are sorted according to adaptation priority and integrated to form a target data processing model set. This target data processing model set provides the foundation for subsequent model status evaluation and data distribution.

[0028] Preferably, the model status evaluation unit constructs a real-time dynamic model status monitoring and evaluation mechanism. Through multi-dimensional indicator collection and comprehensive scoring, it accurately reflects the model's processing capabilities. First, it collects real-time running status data of each data processing model in the target data processing model set. The collected real-time running status data includes key indicators such as the model's current number of connections, CPU utilization, memory usage, average response latency for processing requests in the recent period, request success rate, error return type (such as HTTP 5xx errors, parameter errors, etc.), and cumulative continuous running time. These indicators are obtained in real time through the model running status monitoring interface, and the collection period is set to a reasonable short interval (e.g., once every specific time interval) to ensure the timeliness and accuracy of the data. Subsequently, a multi-dimensional state evaluation model was designed to comprehensively analyze the real-time operating status data of each data processing model. This multi-dimensional state evaluation model includes an indicator weight allocation layer, a normalization processing layer, and a comprehensive scoring layer. The indicator weight allocation layer sets the weights of each operating status indicator according to the scenario requirements. For example, the weights of request success rate and average response latency are higher than those of cumulative continuous runtime, because the former directly affects the user interaction experience. The normalization processing layer converts real-time operating status data of different dimensions into standardized indicator data with a unified value range, eliminating the evaluation interference caused by the difference in dimensions between indicators (such as CPU utilization as a percentage and response latency as a time unit). The comprehensive scoring layer calculates the standardized indicator data based on weighted summation logic to obtain the comprehensive status score of each data processing model. The value range of the comprehensive status score is a reasonable range (such as a certain numerical range). The higher the score, the stronger the current processing capability of the model and the more stable the operating status. Finally, the comprehensive status score of each data processing model is compared with the preset status level classification standard to generate status evaluation data for each data processing model. This status evaluation data includes the comprehensive status score and the corresponding status level (such as excellent, good, medium, poor), and also marks whether the model is currently running abnormally (such as excessive CPU utilization, excessive response latency, excessively low request success rate, etc.), providing a clear basis for subsequent load balancing control.

[0029] Preferably, the load balancing control and data distribution unit module is used to construct a dynamic load balancing mechanism adapted to high-concurrency scenarios. Through differentiated weight allocation and precise data distribution, it improves resource utilization and processing efficiency. The specific technical implementation process is as follows: First, the data processing models in the target data processing model set are classified according to their core data characteristics. Data processing models that adapt to the same data type characteristics and business category characteristics are grouped into the same model group, forming multiple similar model groups. For example, all models adapted to text-based math problem solutions are grouped into one similar model group, and all models adapted to text + image composite essay data are grouped into another similar model group. For each similar model group, the designed dynamic load balancing algorithm is invoked. This algorithm combines the status evaluation data (comprehensive status score, status level, and whether it is abnormal) of each data processing model in the similar model group with the current request queue length of the similar model group to calculate the load allocation weight of each data processing model. The calculation logic for the load allocation weight is: the higher the comprehensive status score and the shorter the current request queue length, the higher the load allocation weight; if a model is marked as running abnormally, its load allocation weight is set to zero, temporarily excluding it from load allocation. Based on the calculated load allocation weights, a specific load allocation scheme is formulated for each similar model group, clarifying the proportion of standardized multimodal data that each data processing model needs to handle. For example, in a similar model group, if model A has a certain load allocation weight and model B has a different weight, then model A will handle that proportion of the total standardized multimodal data for that group, and model B will handle the other proportion. Subsequently, according to the load allocation scheme, the standardized multimodal data is accurately distributed to the corresponding data processing models in the target data processing model set. Each data processing model receives and processes the allocated standardized multimodal data, performing data processing according to its own model expertise (such as semantic analysis for text models and cross-modal feature fusion for multimodal models), generating its own single-source processing result data. Finally, the single-source processing results data output by all data processing models are collected, and the multi-source single-source processing results data are collaboratively integrated through the designed result fusion logic to eliminate data redundancy (such as duplicate answer fragments) and logical conflicts (such as inconsistent output results from different models). The multi-source collaborative output data is combined according to the business scenario requirements to form a unified format of multi-source collaborative output data. This multi-source collaborative output data will be transmitted to the isolation classification and risk compliance audit module for subsequent processing.

[0030] Optionally, the isolation classification and risk compliance audit module includes: a dual-channel isolation processing unit, which performs dual-channel isolation processing on the multi-source collaborative output data based on a dual-channel transmission protocol of inference data channel and result data channel to obtain the inference data and the result data; and an independent identifier configuration unit, which configures inference identifier field and result identifier field for the inference data and the result data to generate inference data and result data with independent identifiers.

[0031] In implementing the above steps, to address the issue of low review accuracy caused by the mixing of reasoning and answers in multimodal processing scenarios, a design of "dual-channel isolation at the protocol layer - refined configuration of identifier fields" is adopted to achieve physical separation and precise differentiation between reasoning data and result data. This module differs from traditional data transmission methods that share channels and have ambiguous identifiers. It embeds the isolation mechanism into the transmission protocol layer, combined with a scenario-based independent identifier system. This ensures that the review focuses solely on the reasoning process, the results are sent out in a controlled manner, and data traceability is improved, providing efficient support for compliance review in learning machine scenarios. Furthermore, this module consists of a dual-channel isolation processing unit and an independent identifier configuration unit. The former is responsible for the physical isolation and transmission of data, while the latter is responsible for the attribute differentiation and identifier of the data. The two units form a coherent "isolation-identification" process, laying the foundation for subsequent risk screening and compliance verification.

[0032] Preferably, the dual-channel isolation processing unit is used to construct a dual-path data transmission mechanism. Through protocol-layer transmission rule definitions, it addresses the issue of mixed reasoning and answer data. First, it receives multi-source collaborative output data from the model scheduling module. This data includes reasoning process information and final result information generated by the model processing multimodal tasks (such as solving math problems and grading essays). Based on the dual-channel transmission protocol, the protocol predefines core parameters for the reasoning data channel and the result data channel, including transmission ports, data format specifications, transmission priorities, and data verification rules. The reasoning data channel transmits information such as the derivation logic and intermediate thinking steps during the model's conclusion generation process. The result data channel transmits the final conclusion information, such as answers and scores, for the learning machine user. The transmission priority of the reasoning data channel is lower than that of the result data channel to ensure the real-time availability of results for the user. Subsequently, content attribute recognition logic is used to parse the multi-source collaborative output data, extracting content features. These features include the data generation purpose (for review and analysis or user reception), data expression form (deductive description or conclusive description), and business scenario information associated with the data (such as math problem-solving steps or essay scoring results). Data is categorized based on content characteristics. Data representing deductive logic and intermediate thought processes, used for review and analysis, is imported into the reasoning data channel. Data representing the final answer and score, used for user reception, is imported into the result data channel. Physical isolation is achieved through parallel transmission via dual channels. During transmission, each channel performs data integrity checks, using checksum comparison to ensure no loss or tampering during data transmission. This results in independently transmitted and complete reasoning and result data. Both types of data carry trace identifiers (Trace-IDs) corresponding to the original standardized multimodal data, ensuring the relevance for subsequent review and result traceability.

[0033] Preferably, the core of the independent identifier configuration unit's module implementation is to construct a scenario-based identifier field system. Through multi-dimensional identifier configuration, it achieves accurate differentiation between inference-type data and result-type data. First, it acquires the inference-type data and result-type data obtained after dual-channel isolation processing. Based on the requirements of multi-modal processing scenarios, it designs a dedicated identifier field system for the two types of data. An inference identifier field is configured for the inference-type data. This inference identifier field includes an inference process identifier, an inference model tracing identifier, and an inference stage identifier. The inference process identifier is used to distinguish the type of inference logic, the inference model tracing identifier is used to record the data processing model information that generated the inference data, and the inference stage identifier is used to mark specific steps in the inference process (such as the initial hypothesis step, the verification step, the conclusion derivation step, etc.). A result identifier field is configured for result-type data. This field includes a result type identifier, a result compliance pre-verification identifier, and a result association inference identifier. The result type identifier distinguishes the presentation format of the result (e.g., text answer, score rating, image annotation result, etc.). The result compliance pre-verification identifier marks whether the result has undergone basic format compliance checks (e.g., whether there are format errors, whether it conforms to the result specifications of the business scenario). The result association inference identifier associates the result with the inference process identifier of the corresponding inference-type data, achieving a one-to-one correspondence between the result and the inference process, facilitating root cause tracing in subsequent audit anomalies. Subsequently, the configured inference identifier field is embedded in the protocol header of the inference-type data, and the result identifier field is embedded in the protocol header of the result-type data. The embedding position and field length are fixed to ensure rapid parsing and recognition by subsequent modules. Finally, the data after embedding the identifier fields is format-integrated to generate inference-type data and result-type data with independent identifiers. The identifier fields of both types of data are stored in association with the data content, ensuring that data types can be quickly distinguished and data sources traced through identifiers during subsequent risk screening and compliance verification.

[0034] Optionally, the data standardization construction module, the model scheduling module, and the isolation audit and compliance processing module have an integrated interface layer structure, which constructs a communication mode with unified access specifications, real-time interaction, and streaming output.

[0035] This application addresses the issues of heterogeneous interfaces, high interaction latency, and limited data transmission modes among modules in multimodal data processing scenarios. It employs an integrated design of "unified access specifications, multimodal interaction adaptation, and streaming output optimization" to construct a unified communication bridge connecting the data standardization construction module, model scheduling module, and isolated audit and compliance processing module. Unlike traditional designs with independent interfaces and fragmented interaction protocols for each module, this approach integrates the communication needs of different modules into a unified interface abstraction. It adapts to synchronous, asynchronous, and streaming transmission scenarios, reducing coupling between modules while ensuring high-concurrency, low-latency multimodal data transmission requirements, providing efficient communication support for text, image, and other multimodal interactions in learning machines. The integrated interface layer structure adopts a layered architecture design, including an interface specification definition layer, a communication mode adaptation layer, and a data transmission optimization layer. These three layers work together to achieve the core functions of unified access, real-time interaction, and streaming output, ensuring consistency, efficiency, and flexibility in data transmission among modules.

[0036] Preferably, the module implementation of the interface specification definition layer is used to build a unified access standard that adapts to multi-module communication. Through standardized protocol formats and field definitions, it solves the problem of incompatibility between traditional module interfaces. The specific technical implementation process is as follows: First, the communication requirements of the data standardization construction module, model scheduling module, and isolation audit and compliance processing module are sorted out. The input and output data formats, interaction frequencies, and error handling requirements between each module are clarified. Based on the Hypertext Transfer Protocol (HTTP), a unified Representational State Transfer (REST) ​​interface specification is designed. This interface specification defines a unified request method (such as get, submit, update, delete), request header format, response code system, and data exchange format (using JavaScript Object Notation, JSON). For core data transmitted between modules (such as standardized multimodal data, target data processing model set information, and multi-source collaborative output data), unified field naming rules, data type constraints, and mandatory field requirements are defined. For example, when transmitting standardized multimodal data, core fields such as modality type identifier, standardized coded data, and business category identifier must be included. Field names uniformly use lowercase underscore naming conventions, and data types are strictly limited to strings, numbers, or booleans to avoid parsing errors caused by inconsistent fields. Simultaneously, a unified error response format is designed, including error code, error description, and error location information. Error codes are categorized and coded according to module affiliation and error type (e.g., errors related to the data standardization module are identified with a specific prefix, and errors related to the model scheduling module are identified with another specific prefix), facilitating rapid problem location by the receiver. These specifications are integrated into an interface specification document and synchronously deployed to all related modules to ensure that each module sends and receives data according to a unified specification, achieving seamless inter-module integration. For example, standardized multimodal data output by the data standardization construction module is encapsulated according to the unified interface specification and submitted to the model scheduling module. The model scheduling module parses the data according to the same specification, requiring no additional format conversion processing.

[0037] Preferably, the core of the communication mode adaptation layer is to build a multi-scenario adaptable interaction mechanism. By dynamically selecting communication protocols, it meets the real-time and flexibility requirements of different modules. In its specific technical implementation, communication mode types are first divided based on the interaction scenario characteristics between modules, including synchronous communication scenarios (such as the data standardization construction module submitting data to the model scheduling module and waiting for an immediate response), asynchronous communication scenarios (such as the model scheduling module pushing multi-source collaborative output data to the isolation audit and compliance processing module without immediate feedback), and streaming communication scenarios (such as the isolation audit and compliance processing module outputting fragmented audit results to the learning machine client). For synchronous communication scenarios, a unified REST interface specification is directly adopted, and data transmission is achieved through HTTP short connections to ensure the immediacy of requests and responses. This is suitable for scenarios with small data volumes and low interaction frequencies (such as the model scheduling module querying data format verification rules from the data standardization construction module). For asynchronous communication scenarios, the WebSocket protocol is introduced to build a bidirectional long-connection communication channel. This protocol supports full-duplex communication between the server and the client (or between modules), eliminating the need for frequent connection establishment, reducing communication latency, and is suitable for scenarios with large data volumes and high real-time requirements (such as the model scheduling module continuously pushing multi-source collaborative output data to the isolation audit and compliance processing module). Identity is verified through a unified handshake protocol during connection establishment to ensure communication security. Heartbeat packets are sent periodically during the connection process to maintain the connection status, and a reconnection mechanism is automatically triggered in case of abnormal disconnection. For streaming communication scenarios, the Server-Sent Events (SSE) protocol is adopted. This protocol supports unidirectional streaming data push from the server to the client (or between modules), suitable for scenarios where data is continuously output in fragments (such as the isolation audit and compliance processing module pushing compliant outgoing data in fragments to the learning machine client). The protocol defines a unified fragment data format, including fragment identifier, sequence number, data content, and end identifier, ensuring that the receiver concatenates data in order. It also supports breakpoint resumption; when transmission is interrupted, the receiver can request the server to continue pushing subsequent fragment data based on the last received sequence number. The communication mode adaptation layer has built-in scene recognition logic, which automatically selects the corresponding communication protocol according to the business scenario of data transmission (such as data type, transmission volume, and real-time requirements). For example, it selects the REST protocol when transmitting a single standardized multimodal data and selects the SSE protocol when transmitting continuous multi-source collaborative output data fragments. Dynamic adaptation of communication mode can be achieved without manual intervention.

[0038] Preferably, the data transmission optimization layer module is implemented to improve the efficiency and reliability of data transmission. Through data compression, transmission priority scheduling, and error recovery mechanisms, it addresses data transmission latency and loss issues in high-concurrency scenarios. The specific technical implementation process is as follows: First, the data transmitted between modules is dynamically compressed. For text-based data (such as text content and error response information in standardized multimodal data), lossless compression algorithms (such as the Deflate algorithm) are used to reduce the data transmission volume. For binary-to-text data such as image Base64 encoded data, efficient text compression algorithms (such as the LZ77 algorithm) are used to balance compression efficiency and decompression speed. Then, a data transmission priority scheduling mechanism is constructed. Priority levels are assigned based on the business importance of the data (e.g., high priority is used for transmitting real-time interactive data from learning machine users, medium priority for transmitting model status information, and low priority for transmitting log data). The transmission queue is sorted by priority, with high-priority data occupying transmission resources first to ensure the real-time transmission of critical data. For example, standardized multimodal data corresponding to essay correction requests submitted by learning machine users is marked as high priority and transmitted to the model scheduling module first, avoiding response delays caused by queue congestion. Simultaneously, a data transmission error recovery mechanism is designed. For data transmitted via the REST protocol, a request-based retransmission mechanism is employed. When the receiver does not receive a response within a preset time or receives an erroneous response (such as an HTTP 5xx error), retransmission is automatically triggered. The number of retransmissions is configured differently based on the business category (e.g., higher priority data is retransmitted more often than lower priority data). For data transmitted via WebSocket and SSE protocols, a data verification and retransmission mechanism is used. Each data fragment carries a checksum (such as Cyclic Redundancy Check, CRC). The receiver verifies data integrity using the checksum. If data loss or corruption is detected, a retransmission request is sent to the sender, specifying the lost fragment identifier. The sender only retransmits the corresponding fragment data, without needing to retransmit all data, thus reducing transmission redundancy. Furthermore, the data transmission optimization layer monitors the network transmission status in real time. When insufficient network bandwidth is detected, the compression level and transmission rate are automatically adjusted to avoid data transmission timeouts caused by network congestion, ensuring relatively stable transmission performance even in complex network environments.

[0039] Optionally, it also includes a model fault tolerance module, specifically including at least one of the following: a configuration strategy management unit, which configures the running configuration parameters of the data processing model differently according to different business categories, the running configuration parameters including identification information, priority, timeout threshold and maximum number of retries; an anomaly detection and automatic recovery unit, which performs real-time status detection on the data processing model matching the running configuration parameters based on the running configuration parameters configured differently by the configuration strategy management unit according to different business categories, synchronously collects running log statistics on the success rate and response latency data of the model and performs dynamic optimization processing, and if an anomaly is detected in the data processing model, performs self-recovery processing on the data processing model with an anomaly through a service anomaly status reset mechanism, and synchronously feeds back the self-recovery result and anomaly status data to the model switching unit; and a model switching unit, which receives the self-recovery result and anomaly status data fed back by the anomaly detection and automatic recovery unit, and, in combination with the running configuration parameters configured differently by the configuration strategy management unit according to different business categories, triggers a backup data processing model or backup multimodal data processing node adapted to the currently anomaly data processing model.

[0040] The model fault tolerance module of this application addresses the differentiated requirements for the stability of data processing model operation in different business categories (such as math problem solving and essay grading) under multimodal data management scenarios. Through the technical design of "business-adaptive parameter configuration - precise status monitoring and self-recovery - on-demand backup resource switching", it realizes the fully automated handling of data processing model anomalies.

[0041] Specifically, in this application, the core implementation of the configuration strategy management unit is to construct a mapping and association system between business categories and data processing model operation configuration parameters, thereby achieving precise implementation of differentiated configurations based on business categories. Preferably, in a scenario, when the configuration strategy management unit is specifically implemented, it first performs business category feature parsing to extract the core operational requirement features of each business category (e.g., math problem solving business, essay correction business). These core operational requirement features include the multimodal data type combinations corresponding to the business (plain text, text + image), the real-time requirement level of data processing, and the reliability requirement level of data output. Based on the extracted core operational requirement features of the business categories, a business category-operation configuration parameter mapping table is constructed. In the runtime configuration parameter mapping table, each record contains five related fields: business category identifier, data processing model identifier, priority parameter, timeout threshold parameter, and maximum retry count parameter. The business category identifier must match the business category identifier in the standardized multimodal data output by the data standardization construction module to ensure accurate matching of parameter configuration with subsequent data processing models. For each business category's runtime configuration parameters, a parameter rationality check is performed. The check logic combines the runtime logs of the historical data processing models for that business category to determine whether the configured timeout threshold parameter is within a reasonable range of normal model latency for that business category, whether the maximum retry count parameter is compatible with the concurrency of data processing for that business category, and whether the priority parameter conforms to the business importance ranking (e.g., models for core businesses have higher priority than those for ordinary businesses). After successful check, the business category-running configuration parameter mapping table is stored in a distributed parameter library, and an incremental synchronization mechanism for parameter updates is established to ensure that parameter modifications are synchronized in real-time to the anomaly detection and automatic recovery unit and the model switching unit, providing a unified parameter basis for subsequent status monitoring and switching decisions.

[0042] Specifically, in this application, the anomaly detection and automatic recovery unit is based on the business category-run configuration parameter mapping table output by the configuration strategy management unit to achieve accurate monitoring, anomaly judgment, and self-recovery handling of the data processing model's running status. Preferably, in the specific technical implementation of the anomaly detection and automatic recovery unit, dynamic matching of running configuration parameters is first performed. Based on the identification information of the current data processing model to be monitored, the corresponding business category and associated priority parameters, timeout threshold parameters, and maximum retry count parameters are matched from the business category-run configuration parameter mapping table. The matched running configuration parameters are then bound to the data processing model to form a model-parameter binding relationship. Based on the model-parameter binding relationship, a real-time status monitoring thread is started. This thread collects the running status data of the data processing model according to a preset monitoring cycle. The running status data includes the current processing task queue length of the model, the processing time of each task, and the task processing result (success / failure). The system resources (memory, CPU) used by the model process are monitored (failure), and the running logs of the data processing model are collected synchronously. The running logs are then parsed in a structured manner to extract task identifiers, processing start times, processing end times, and error code information, generating structured log records. Based on the extracted structured log records, the success rate and response latency statistics of the model are calculated. The success rate is the ratio of the number of successfully processed tasks to the total number of tasks within a unit of time, and the response latency is the average processing time of all tasks within a unit of time. Combined with the matched timeout threshold parameter and maximum retry count parameter, anomaly detection is performed. The anomaly detection logic includes: when the processing time of a task exceeds the timeout threshold parameter, it is determined to be a timeout anomaly; when the model's success rate within a unit of time is lower than the preset success rate threshold (this threshold is set based on the core operational requirements of the business category), it is determined to be successful. Power anomaly; if the number of retries for the same task reaches the maximum number of retries and still fails to process successfully, it is judged as a retry limit exceedance anomaly; if any of the above anomalies are detected, the service anomaly status reset mechanism is initiated. This mechanism first performs a model lightweight reset, that is, releases the failed tasks currently being processed by the model, clears the cached data of the model process, and reinitializes the model's processing interface; after the lightweight reset is completed, the model's recovery status data is monitored. If the recovery status data shows that the model can normally receive and process new tasks, a self-recovery success result is generated; if the model still cannot run normally after the lightweight reset, the model process is restarted. After the restart is completed, the recovery status data is monitored again, and the corresponding self-recovery result (success / failure) is generated; the generated self-recovery result and the corresponding anomaly status data (including anomaly type, anomaly occurrence time, and the business category identifier associated with the anomaly task) are encapsulated and transmitted to the model switching unit.

[0043] Specifically, in this application, the model switching unit receives the self-recovery result and abnormal status data output by the anomaly detection and automatic recovery unit, and combines this with the operation configuration parameters of the configuration strategy management unit to achieve precise switching of backup resources. Preferably, in one scenario, when the model switching unit is specifically implemented, it first receives the self-recovery result and abnormal status data, parses the abnormal status data, and extracts the anomaly type, abnormal business category identifier, and abnormal data processing model identifier; it then determines whether the self-recovery result is a failure. If the self-recovery result is a success, the subsequent switching process is terminated; if the self-recovery result is a failure, based on the abnormal business category identifier, it starts from the business category... The priority parameters and associated backup resource configuration information corresponding to the business category are matched against the configuration parameter mapping table (the backup resource configuration information is pre-installed in the configuration policy management unit, including a list of backup data processing models and a list of backup multimodal data processing nodes adapted to the business category); based on the matched priority parameters, the backup data processing model list is sorted, with the backup data processing model with higher priority parameters listed first; real-time running status data (processing task queue length, current resource utilization) of each backup data processing model in the sorted list is collected, and the available load capacity of each backup data processing model is calculated. The available load capacity is the maximum processing capacity of the backup data processing model and the current processing task queue. The length difference is considered; the backup data processing model with the largest available load capacity is selected as the target backup model; if the available load capacity of all models in the backup data processing model list cannot meet the task processing requirements of the current abnormal model, then the node associated with the business category is selected from the backup multimodal data processing node list as the target backup node; a switching instruction is generated, which includes the abnormal data processing model identifier, the target backup model identifier (or the target backup node identifier), and the list of tasks to be migrated; the switching instruction is sent to the model scheduling module, which migrates the tasks to be processed by the original abnormal data processing model to the target backup model (or the target backup node), completing the switching of backup resources and ensuring the continuous progress of multimodal data processing tasks for the corresponding business category.

[0044] Optionally, it also includes a full-link monitoring module, which comprises at least one of a distributed tracing unit, a performance data acquisition unit, a multi-dimensional indicator statistics unit, and a data visualization unit. The distributed tracing unit embeds a unique tracking identifier into the entire process of the standardized multimodal data processing based on the full-link tracing mechanism to construct a full-process tracing link covering the data standardization construction module, the standardization construction module, and the isolation classification and risk compliance review module. The performance data acquisition unit collects request processing runtime data during the data request processing process based on the unique tracking identifier of the distributed tracing unit. The request processing runtime data includes response time and call success rate. The system receives the request processing operation data and error type data, and synchronously transmits the request processing operation data, carrying a unique identifier, to the multi-dimensional indicator statistics unit. The multi-dimensional indicator statistics unit, based on the unique tracking identifier of the distributed tracking unit, combines the data processing model dimension and business category dimension to calculate multi-dimensional business operation statistics data, including call volume, average response latency, percentage of degradation trigger times, and percentage of audit rule hit times. The data visualization unit receives the request processing operation data and the multi-dimensional business operation statistics data, and displays the request processing operation data and the multi-dimensional business operation statistics data in a structured dashboard format.

[0045] This application's end-to-end monitoring module addresses the core needs of end-to-end traceability and operational status awareness in multimodal data management scenarios. Through a progressive technical design—"unique tracking identifier embedded throughout the entire process - precise collection of associated data - multi-dimensional indicator statistics - structured visualization"—it constructs an end-to-end monitoring system covering core aspects such as data standardization, model scheduling, and risk and compliance auditing. The core of this embodiment lies in establishing a deep binding mechanism between the unique tracking identifier and the entire multimodal data processing process. Relying on identifier association, it achieves the serial collection and statistics of cross-module operational data. Unlike traditional distributed monitoring solutions, it enables end-to-end status awareness of the entire multimodal data processing process, providing accurate data support for fault location and system optimization, and adapting to the needs of multimodal data processing scenarios involving multiple business categories and multi-model collaboration.

[0046] Specifically, in this application, the core implementation of the distributed tracing unit is to construct a mechanism for generating, embedding, and transmitting unique tracking identifiers throughout the entire process, ensuring the traceability of each stage of multimodal data processing. Preferably, in one scenario, the distributed tracing unit first generates a unique tracking identifier using a combination of timestamp, business category identifier, and random sequence. The timestamp is the precise time of identifier generation (accurate to milliseconds), the business category identifier is extracted from the business information associated with the initial multimodal data received from the data standardization construction module (maintaining consistency with the business category identifier in the data standardization construction module, such as a specific business category identifier corresponding to a math problem solution), and the random sequence is a character combination of a preset length. The unique tracking identifier generated through this combination rule ensures global uniqueness and avoids identifier conflicts between different data processing flows. After generating the unique tracking identifier, it is embedded into the header field of the structured multimodal data output by the data standardization construction module, forming structured multimodal data with a unique tracking identifier. During the embedding process, the integrity of the original fields of the structured multimodal data is maintained. In the communication mode of the integrated interface layer structure, when structured multimodal data with unique tracking identifiers flows to subsequent stages such as the model scheduling module and the isolation classification and risk compliance review module, the identifier synchronous transmission process is performed. That is, after receiving data, each module synchronizes the unique tracking identifier in the header of its own output data to ensure that the unique tracking identifier is carried throughout the entire data flow. At the same time, a unique tracking identifier field is added to the processing log of each module, so that the operation log of each module can be accurately associated with the corresponding data processing flow. Finally, a full-process tracking link covering the data standardization construction module, model scheduling module, isolation classification and risk compliance review module is constructed. This full-process tracking link can connect the operation information of each stage of data processing through the unique tracking identifier.

[0047] Specifically, in this application, the performance data acquisition unit uses the unique tracking identifier generated by the distributed tracing unit as the core association basis to achieve accurate acquisition and associated transmission of request processing operation data throughout the multimodal data processing process. Preferably, in the specific technical implementation of the performance data acquisition unit, the acquisition rule configuration is first executed. Based on the core links of multimodal data processing (data standardization construction, model scheduling, and risk compliance review), the acquisition items and acquisition cycle for each link are configured. The acquisition items include response time, call success rate, and error type. The acquisition cycle is configured to different durations according to the processing time characteristics of each link (e.g., the data standardization construction link has a shorter processing time, so the acquisition cycle is configured to a shorter duration; the model scheduling link has a relatively longer processing time, so the acquisition cycle is configured to a longer duration). Based on the configured acquisition rules, the acquisition thread is started, and the processing data of each link is associated through the unique tracking identifier. The acquisition logic for response time is to record the difference between the time data enters the current link and the time data leaves the current link to obtain the response time of the current link; the acquisition logic for call success rate is to uniformly... Within a given collection period, the ratio of the amount of data successfully processed in the current stage to the total amount of data processed yields the success rate of the current stage's call. The error type collection logic involves parsing error records with unique tracking identifiers in the current stage's operation log, extracting error codes and descriptions, and classifying the error records according to preset error type classification rules (such as format errors, model call errors, and data transmission errors) to obtain error type data. After collecting request processing operation data such as response time, call success rate, and error type, a corresponding unique tracking identifier is added to each request processing operation data entry, forming request processing operation data with unique tracking identifiers. This request processing operation data with unique tracking identifiers is then synchronously transmitted to the multi-dimensional indicator statistical unit through a real-time interactive channel of an integrated interface layer structure, ensuring accurate correlation between the collected data and subsequent statistical stages.

[0048] Specifically, in this application, the multi-dimensional indicator statistics unit is based on the request processing operation data with unique tracking identifiers transmitted by the performance data acquisition unit, and combines the data processing model dimension and business category dimension to accurately generate multi-dimensional business operation statistics data. Preferably, in a scenario, when the multi-dimensional indicator statistics unit is specifically implemented, it first performs the reception and parsing of request processing operation data with unique tracking identifiers, extracting the unique tracking identifier, response time, call success rate, error type, and corresponding data processing model identifier and business category identifier from the data (the data processing model identifier and business category identifier are extracted from the structured multimodal data associated with the unique tracking identifier); based on the extracted data processing model identifier and business category identifier, a two-dimensional statistical dimension index is constructed, where one dimension is the data processing model dimension (distinguished by the data processing model identifier), and the other dimension is the business category dimension (distinguished by the business category identifier); for the data processing model dimension, the call volume, average response latency, percentage of downgrade triggers, and percentage of audit rule hits corresponding to each data processing model are statistically analyzed, where the call volume is the total amount of data with unique tracking identifiers processed by the model per unit time, the average response latency is the average of all response times corresponding to the model, and the percentage of downgrade triggers is the average of the data processing model's response time. The ratio of the number of times the model triggers degradation processing to the total number of calls (the number of degradation triggers is extracted from the degradation records with unique tracking identifiers transmitted by the model fault tolerance module). The percentage of audit rule hits is the ratio of the number of times the data processed by the model hits the rules after being reviewed by the isolation classification and risk compliance audit module to the total number of calls (the number of audit hits is extracted from the audit records with unique tracking identifiers transmitted by the isolation classification and risk compliance audit module). For the business category dimension, the call volume, average response latency, percentage of degradation triggers, and percentage of audit rule hits corresponding to each business category are statistically analyzed. The statistical logic is consistent with the data processing model dimension, only the statistical dimension is replaced with the business category identifier. After the statistics are completed, multi-dimensional business operation statistics are generated. Each multi-dimensional business operation statistics are associated with the corresponding data processing model identifier, business category identifier, and unique tracking identifier fragment (the timestamp part of the unique tracking identifier is extracted for time dimension correlation analysis) to ensure the traceability and relevance of the statistics.

[0049] Specifically, in this application, the core implementation of the data visualization unit is to receive multi-dimensional business operation statistics output by the multi-dimensional indicator statistics unit and request processing operation data output by the performance data acquisition unit, and to realize the hierarchical display of data through a structured dashboard. Preferably, in the specific technical implementation of the data visualization unit, the data reception and integration are performed first. This involves receiving multi-dimensional business operation statistics output by the multi-dimensional indicator statistics unit and request processing operation data output by the performance data acquisition unit. Based on unique tracking identifiers and associated data processing model identifiers and business category identifiers, data from the same data processing process, the same data processing model, and the same business category are integrated to form an integrated structured data set. Based on this integrated structured data set, a multi-level structured dashboard is constructed. The first level is a global overview dashboard, displaying the total call volume, overall average response latency, overall call success rate, and the percentage of various error types for the entire system, presented in numerical, line, and pie chart formats. The second level is a business category dimension dashboard, categorized by business category identifier, displaying the call volume, average response latency, percentage of degradation triggers, and percentage of audit rule hits for each business category, using bar charts to compare the operational indicators of different business categories. The third level is a data processing model dimension dashboard, categorized by data processing model identifier, displaying the call volume, average response latency, percentage of degradation triggers, and percentage of audit rule hits for each model, using tables and... The line chart presents the model's operational trend; the fourth layer is a detailed traceability dashboard, which retrieves the response time, processing results, and error messages (if any) of each stage of the corresponding data processing flow through a unique tracking identifier, enabling accurate fault tracing; the structured dashboard supports interactive search functionality, allowing users to quickly locate the corresponding operational data by inputting a unique tracking identifier, business category identifier, or data processing model identifier. The integrated structured data set is then rendered according to the above dashboard hierarchy and display format, completing the structured dashboard display of request processing operational data and multi-dimensional business operational statistics, providing operations and maintenance personnel with an intuitive perception of the entire operational status.

[0050] Optionally, it also includes a multi-layer risk control module, which specifically has at least one of a dual-review unit, a dynamic downgrade handling unit, and a streaming gating review unit: The dual-review unit includes an input-end risk identification unit and an output-end compliance verification unit. The input-end risk identification unit identifies pre-risk information such as risk identification and risk level for text-sensitive information and image-sensitive risks in the initial multimodal data. The output-end compliance verification unit performs secondary verification on the compliance data to be sent out. The dynamic downgrade handling unit collects the pre-risk information identification results, including risk identification and risk level, and combines the risk attributes of the data processing model to differentiate handling strategies based on risk level, so as to transmit the risk identification and handling strategies to the business layer and the streaming gating review unit. The streaming gating review unit receives the risk identification and handling strategies transmitted by the dynamic downgrade handling unit, performs streaming processing and real-time risk matching on the standardized multimodal data using a segmented review mode based on the handling strategies, and performs differentiated processing according to the matching results.

[0051] This application's multi-layered risk control module addresses the differentiated compliance needs of various business categories (such as math problem solving and essay grading) in multimodal data management scenarios. Through a hierarchical design of "pre-emptive risk identification - dynamic strategy adaptation - streaming precision review," it constructs a risk prevention and control system covering the entire process of data input, processing, and output. The core of this application lies in establishing a dynamic binding mechanism between risk identification and handling strategies. Relying on the modal characteristics and business attributes of multimodal data, it achieves precise and real-time risk prevention and control. Unlike traditional, singular, and generalized risk control solutions, it can adapt to the complex risk characteristics of multimodal data (text, images, and text + images), ensuring real-time processing efficiency while improving the targeting and accuracy of compliance verification, meeting the compliance output requirements of multiple business scenarios.

[0052] Specifically, in this application, the core implementation of the dual-audit unit is to construct a hierarchical audit mechanism for the input and output ends. Through progressive processing of pre-risk identification and post-compliance verification, it achieves end-to-end risk filtering of multimodal initial data. Preferably, in one scenario, when the dual-audit unit is specifically implemented, the input-end risk identification unit first receives the multimodal initial data and sorts it according to data modality type (text, image, text+). The image data undergoes modal splitting to obtain single-modal text data, single-modal image data, or mixed-modal data. For single-modal text data, a designed text-sensitive feature matching algorithm is used. This algorithm first performs word segmentation on the text data, extracts the word segmentation results, and then matches the word segmentation results with a preset text-sensitive feature library. During the matching process, the semantic similarity between the word segmentation results and the sensitive features is calculated (semantic similarity is calculated through word vector space distance). When the semantic similarity reaches a preset threshold, it is marked as sensitive text, and a corresponding text risk label and risk level are generated (risk levels are divided into different levels according to the semantic similarity). For single-modal image data, a designed image sensitive region detection model is used. This model includes an image preprocessing layer, a feature extraction layer, and a risk classification layer. The image preprocessing layer performs size normalization and pixel standardization on the image data to obtain standardized image data. The feature extraction layer extracts texture and contour features from the standardized image data to obtain image risk features. The risk classification layer... Image risk features are compared with a pre-set image sensitive feature library to output image risk identifiers and risk levels. For mixed-modal data, the above-mentioned text sensitive feature matching and image sensitive area detection processing are performed simultaneously to integrate and obtain comprehensive risk identifiers and risk levels, forming a preliminary risk information identification result. The output compliance verification unit receives the compliance data to be sent out after processing by the isolation classification and risk compliance review module, extracts the result data (carrying an independent result identifier field) from the compliance data to be sent out, and performs a secondary matching verification with the compliance feature library of the corresponding business category (matched according to the business category identifier, such as the compliance feature library of the mathematics field for math problem solving business, and the compliance feature library of the essay field for essay correction business). The matching content includes the standardization of the expression of the result data, the accuracy of the information, and compliance. If no risk is found in the secondary matching verification, the outgoing data that meets the compliance requirements is output. If a risk is found, the corresponding result data is returned to the isolation classification and risk compliance review module for reprocessing to ensure the compliance of the outgoing data.

[0053] Specifically, in this application, the dynamic degradation handling unit, based on the pre-risk information identification results output by the dual review unit, combines the risk attributes of the data processing model to achieve differentiated matching and precise push of handling strategies. Preferably, in the specific technical implementation of the dynamic degradation handling unit, the receiving and parsing of the pre-risk information identification results are performed first, extracting the risk identifier, risk level, and corresponding multimodal initial data modality type and business category identifier from the results; simultaneously collecting the risk attribute information of the data processing model, which includes the business categories supported by the model, historical risk handling records, and risk bearing thresholds, the risk attribute information of the data processing model is extracted from the model information registry of the model scheduling module; based on the extracted risk level and the risk attributes of the data processing model, a risk level-handling strategy mapping rule is constructed, which divides different handling strategies according to the risk level. The handling strategy corresponding to the low risk level is to retain the risk identifier and continue normal processing, the handling strategy corresponding to the medium risk level is to enable the security template to replace the risk content before processing, and the handling strategy corresponding to the high risk level is to block data transmission and trigger an alarm; according to the risk level-handling strategy mapping rule, the corresponding handling strategy is matched for the currently parsed risk identifier, generating risk identifier-handling strategy associated data; the risk identifier-handling strategy is then linked. The data associated with the handling strategy is transmitted to the business layer and the streaming gating audit unit, respectively. The data transmitted to the business layer is used to optimize the risk control strategy corresponding to the business category, and the data transmitted to the streaming gating audit unit is used for strategy adaptation in subsequent streaming audits.

[0054] Specifically, in this application, the streaming gating audit unit uses the risk identifier-disposal strategy association data output by the dynamic degradation disposal unit as its core basis, and combines the streaming processing characteristics of multimodal data to achieve real-time and accurate risk auditing of standardized multimodal data. Preferably, in a scenario, when the streaming gating audit unit is specifically implemented, it first receives the risk identifier-disposal strategy association data transmitted by the dynamic degradation disposal unit, and parses it to obtain the risk identifier and the appropriate disposal strategy corresponding to the current data to be audited; based on the streaming output communication mode of the integrated interface layer structure, it receives the standardized multimodal data output by the data standardization construction module, and splits the standardized multimodal data into multiple data fragments according to the preset fragmentation rules (fragmentation rules are divided according to data size or data semantic integrity) to obtain streaming data fragments; for each streaming data fragment, it extracts the modality type and business category identifier of the data fragment, and performs accurate risk matching processing in combination with the parsed risk identifier, that is, text data fragments use the above-mentioned text sensitive feature matching algorithm, image data fragments use the above-mentioned image sensitive area detection model, and mixed data fragments perform both processing simultaneously; based on the risk matching result and the corresponding The handling strategy is differentiated. If a risk identifier is matched and the corresponding handling strategy is to retain the risk identifier and process it normally, the risk identifier is marked and transmitted to the subsequent module. If the handling strategy is to enable security template replacement, the security template of the corresponding business category (such as the formula specification template for math problem solving business and the expression specification template for essay correction business) is called to replace the risk content and obtain a security data fragment. If the handling strategy is to block data transmission, the transmission of the current data fragment is stopped, a risk alarm message is generated and fed back to the dynamic degradation handling unit. After all streaming data fragments are processed, fragment integration processing is performed to obtain integrated standardized multimodal data. The integrated standardized multimodal data is transmitted to the model scheduling module to ensure the controllability of risks in subsequent data processing.

[0055] like Figure 2 The present application also provides a multimodal data management method, including the following steps: performing unified encapsulation and format standardization verification on initial multimodal data to generate standardized multimodal data; determining a data processing model adapted to the standardized multimodal data based on the data type and business category of the standardized multimodal data, and performing load balancing processing on the data processing models adapted to the same data type to generate multi-source collaborative output data; splitting the multi-source collaborative output data based on a dual-channel isolation mechanism to obtain inference data and result data, and assigning independent identifiers to them for differentiation; performing risk screening on the inference data and compliance verification on the result data based on the independent identifiers and a preset risk feature library, to output outgoing data that has been risk-filtered and meets compliance requirements.

[0056] This application addresses the issues of fragmented data processing, insufficient compliance, and poor real-time performance in multimodal interaction scenarios. It employs a four-step closed-loop design—standardized encapsulation, intelligent scheduling, isolation labeling, and risk management—to achieve efficient and compliant end-to-end processing of multimodal data from input to output. This method differs from traditional step-by-step processing approaches by deeply integrating data standardization with business scenarios, dynamically adapting scheduling strategies to model states, and precisely linking isolation mechanisms with risk audits. This not only solves the pain points of messy multimodal data formats and low processing efficiency but also ensures content compliance and real-time interaction in learning machine scenarios, adapting to the complex processing needs of multimodal data such as text and images.

[0057] Step 1: Unified encapsulation and processing of multimodal data and standardization and validation of format; Preferably, the system first receives multimodal initial data input by the user. This data covers various forms such as math problem text, essay content, solution diagrams, and knowledge point illustrations. Based on modality recognition logic, the system extracts the morphological features of the data to determine whether it is text-based, image-based, or a text + image composite, generating a corresponding modality type identifier (MessageType). For different modality types, scenario-based standardized encoding is performed. Text-based data is converted using NFKC character standardization encoding rules to eliminate character differences caused by different input devices. Image-based data first removes Exchangeable Image File Format (EXIF) information, then converts it to the sRGB color space, adjusts the size according to a preset ratio, saves it as JPEG or PNG format with reasonable compression quality, and finally converts it to Base64 encoding. For text + image composite data, after completing the standardized encoding of both text and image, a modality association identifier is generated to record the semantic correspondence between the two (e.g., the attribution association between essay text and accompanying illustrations). Subsequently, a business category identifier (BusinessID) is matched according to the learning scenario corresponding to the data, including math problem solutions, essay correction, etc., with each identifier corresponding to a specific processing strategy. Following the field rules of "Modal Type Identifier - Standardized Encoded Data - Modal Association Identifier (Composite Data Only) - Business Category Identifier", the data is encapsulated into structured multimodal data (MultiModalMessage). The structured multimodal data undergoes format standardization validation, verifying field completeness, encoding validity, and structural regularity. For example, it checks whether text encoding conforms to the NFKC standard, whether image Base64 encoding is decodeable, and whether the business category identifier is valid. Validated data is output as standardized multimodal data; otherwise, an anomaly report is generated and sent back to the user.

[0058] Step 2: Determining the appropriate data processing model and handling load balancing; Preferably, core features, including modality type identifiers and business category identifiers, are first extracted from standardized multimodal data to form data processing requirement tags. A pre-defined model mapping rule base is invoked. This rule base adopts a two-dimensional association structure of "modality type + business category" with the data processing model. For example, pure text math problems correspond to algebraic problem-solving models in the text model pool, and math problems containing geometric figures correspond to graphic-text problem-solving models in the multimodal model pool. All suitable data processing models are selected through feature matching to form a target data processing model set. Real-time running status data of each model in the target data processing model set is collected, including indicators such as CPU utilization, memory usage, response latency, and request success rate. A multi-dimensional evaluation logic is used to generate status evaluation data, intuitively reflecting the model's current processing capacity. Based on the status evaluation data, load balancing is performed on models with the same data features. A dynamic weight allocation algorithm distributes standardized multimodal data proportionally to each model, avoiding excessive load on any single model. After receiving the data, each model performs targeted processing. The text model focuses on text semantic analysis and logical deduction, while the multimodal model realizes cross-modal feature fusion and collaborative inference between text and images. After processing, each model outputs its own output data. All model output data are collected, and the results are integrated, deduplicated, and conflict-eliminating through result fusion logic to generate multi-source collaborative output data in a unified format.

[0059] Step 3: Dual-channel isolation processing and independent identification configuration; Preferably, the multi-source collaborative output data is isolated based on the designed dual-channel transmission protocol. First, the data content attributes are parsed to distinguish between the model-generated reasoning trace and the final result (answer). The reasoning trace includes the model's thought process and logical deductions, while the final result provides user-facing answers, ratings, and other conclusive information. The reasoning trace data is imported into the reasoning data channel, and the final result data is imported into the result data channel, achieving physical isolation transmission and ensuring that the review process only targets the reasoning data, without affecting the real-time output of the results. Independent identification fields are configured for the two types of data. The identification field for reasoning data includes reasoning model traceability information and reasoning stage markers, used to trace the source and stages of the reasoning process. The identification field for result data includes result type, compliance pre-verification markers, and associated reasoning markers, used to associate the corresponding reasoning process, ensuring result traceability. The identified reasoning data and result data are then transmitted to the risk screening and compliance verification stages respectively. The reasoning data is used by the review engine to analyze risks, while the result data awaits compliance verification before being released.

[0060] Step Four: Risk Screening and Compliance Verification; Preferably, a pre-defined risk feature library is first invoked. This feature library is built based on the learning machine scenario and includes a sensitive word library suitable for teenagers, illegal image features, and risk expression patterns, divided into three levels: low, medium, and high, according to the severity of the risk. For reasoning data with independent identifiers, a risk screening logic combining dictionary matching, regular expression detection, and lightweight semantic analysis is adopted. Dictionary matching is used to identify sensitive words, regular expression detection is used to match illegal information in specific formats (such as contact information and illegal URLs), and lightweight semantic analysis is used to determine the illegal intent related to the context, such as identifying sensitive logic or illegal derivation directions contained in the reasoning process. After screening, the risk level and risk type of the reasoning data are marked. Based on the risk screening results of the reasoning data, compliance verification is performed on the result data with independent identifiers to verify whether the result data is consistent with the approved reasoning process and whether there is a logical break due to risk substitution. At the same time, a secondary verification is performed to check whether there are hidden risks (such as implicitly expressed sensitive information). For data that passes verification, it is marked as "compliant and passed" and pushed to the client in streaming mode. For data that fails verification, it is processed differently according to the risk level. Low-risk data is replaced with a security template to replace the illegal parts, while medium- and high-risk data is directly blocked from output and a compliance prompt is returned. Finally, the data that has been filtered for risk and meets the compliance requirements is output to the outside world.

[0061] like Figure 3 Furthermore, this application embodiment also provides an electronic device, comprising: a memory for storing executable program code; and a processor for calling and running the executable program code from the memory, thereby enabling the electronic device to implement the aforementioned multimodal data management method. The electronic device includes, but is not limited to, laptops, smartphones, tablets, smartwatches, etc.

[0062] This application also provides a computer storage medium storing at least one instruction or at least one program, which is loaded and executed by a processor to implement the above-described multimodal data management method.

[0063] This application also provides a computer program product or computer program, which includes computer instructions that, when executed by a processor, implement the above-described multimodal data management method.

[0064] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A multimodal data management system, characterized in that, include: The data standardization construction module is used to perform unified encapsulation and format standardization verification of the initial multimodal data to generate standardized multimodal data. The model scheduling module determines a data processing model that is compatible with the standardized multimodal data based on the data characteristics of the standardized multimodal data, and performs load balancing processing on the data processing models that are compatible with the same data characteristics to generate multi-source collaborative output data. The isolation classification and risk compliance audit module isolates and classifies the multi-source collaborative output data, and performs risk screening and compliance verification on the isolated and classified multi-source collaborative output data in combination with a preset risk feature library, so as to output outgoing data that meets compliance requirements.

2. The multimodal data management system according to claim 1, characterized in that, The data standardization construction module includes: The multimodal data unified encapsulation unit encapsulates the initial multimodal data based on the field composition rules of data type, multimodal data standardized coding and business category identifier to form structured multimodal data; The format standardization verification unit performs format standardization verification on the structured multimodal data and outputs standardized multimodal data that conforms to the format standardization verification.

3. The multimodal data management system according to claim 1, characterized in that, The model scheduling module includes: The data feature acquisition and model mapping unit determines the target data processing model set based on the preset data processing model and the mapping rules of the data type and business category of the standardized multimodal data; The model status evaluation unit generates status evaluation data for each data processing model in the target data processing model set based on the real-time running status data of each data processing model in the target data processing model set. The load balancing control and data distribution unit performs load balancing processing on data processing models that are adapted to the same data characteristics based on the status evaluation data of each data processing model, and distributes the standardized multimodal data to the target data processing model set to generate multi-source collaborative output data.

4. The multimodal data management system according to claim 1, characterized in that, The segregation classification and risk compliance review module includes: The dual-channel isolation processing unit performs dual-channel isolation processing on the multi-source collaborative output data based on a dual-channel transmission protocol for inference data channels and result data channels to obtain the inference data and the result data. An independent identifier configuration unit configures an inference identifier field and a result identifier field for the inference data and the result data to generate inference data and result data with independent identifiers.

5. A multimodal data management system according to claim 1, characterized in that, The data standardization construction module, the model scheduling module, and the isolation audit and compliance processing module have an integrated interface layer structure, which establishes a communication mode with unified access specifications, real-time interaction, and streaming output.

6. A multimodal data management system according to claim 1, characterized in that, It also includes a model fault tolerance module, specifically including at least one of the following: The configuration strategy management unit configures the running configuration parameters of the data processing model differently according to different business categories. The running configuration parameters include identification information, priority, timeout threshold and maximum number of retries. The anomaly detection and automatic recovery unit, based on the operation configuration parameters configured differently by the configuration strategy management unit according to different business categories, performs real-time status detection on the data processing model that matches the operation configuration parameters, synchronously collects operation log statistics on the success rate and response latency data of the model and performs dynamic optimization processing. If an anomaly is detected in the data processing model, the unit performs self-recovery processing on the data processing model with an anomaly through the service anomaly status reset mechanism, and synchronously feeds back the self-recovery results and anomaly status data to the model switching unit. The model switching unit receives the self-recovery results and abnormal status data from the anomaly detection and automatic recovery unit. Combined with the operation configuration parameters configured differently according to different business categories by the configuration strategy management unit, it triggers a backup data processing model or backup multimodal data processing node that is adapted to the currently abnormal data processing model.

7. A multimodal data management system according to any one of claims 1-6, characterized in that, It also includes a full-link monitoring module, which comprises at least one of a distributed tracing unit, a performance data acquisition unit, a multi-dimensional indicator statistics unit, and a data visualization unit, wherein: The distributed tracking unit embeds a unique tracking identifier into the entire process of the standardized multimodal data processing based on the end-to-end tracking mechanism, so as to build an end-to-end tracking link covering the data standardization construction module, the standardization construction module, and the isolation classification and risk compliance review module; The performance data acquisition unit collects request processing operation data during the data request processing process based on the unique tracking identifier of the distributed tracing unit. The request processing operation data includes response time, call success rate and error type data, and synchronously transmits the request processing operation data carrying the unique identifier to the multi-dimensional indicator statistics unit. The multi-dimensional indicator statistical unit, based on the unique tracking identifier of the distributed tracking unit, combines the data processing model dimension and the business category dimension to respectively statistically analyze multi-dimensional business operation statistics. The multi-dimensional business operation statistics include data such as call volume, average response latency, percentage of downgrade trigger times, and percentage of audit rule hit times. The data visualization unit receives the request processing operation data and the multi-dimensional business operation statistics, and displays the request processing operation data and the multi-dimensional business operation statistics in a structured dashboard format.

8. A multimodal data management system according to claim 1, characterized in that, It also includes a multi-layer risk control module, which specifically has at least one of the following: a dual review unit, a dynamic downgrade processing unit, and a streaming gating review unit: The dual audit unit includes an input-end risk identification unit and an output-end compliance verification unit. The input-end risk identification unit identifies the risk information and risk level of text-sensitive information and image-sensitive risks in the multimodal initial data. The output-end compliance verification unit performs a second verification on the compliance data to be sent out. The dynamic downgrade handling unit collects the pre-risk information identification results, including risk identifiers and risk levels, and combines the risk attributes of the data processing model to match handling strategies according to risk level differences, so as to transmit the risk identifiers and handling strategies to the business layer and the streaming gating review unit. The streaming gating review unit receives risk identifiers and handling strategies transmitted by the dynamic downgrade handling unit. Based on the handling strategies, it performs streaming processing and real-time risk matching on the standardized multimodal data using a segmented review mode, and performs differentiated processing based on the matching results.

9. A multimodal data management method, characterized in that, Includes the following steps: The initial multimodal data is subjected to unified encapsulation and format standardization verification to generate standardized multimodal data. Based on the data type and business category of the standardized multimodal data, a data processing model adapted to the standardized multimodal data is determined, and load balancing processing is performed on the data processing models adapted to the same data type to generate multi-source collaborative output data; The multi-source collaborative output data is split based on a dual-channel isolation mechanism to obtain inference data and result data, and each is assigned an independent identifier to distinguish them. Based on the independent identifier, risk screening is performed on the inference data and compliance verification is performed on the result data in combination with a preset risk feature library, so as to output outgoing data that has been risk-filtered and meets compliance requirements.

10. An electronic device, characterized in that, The electronic device includes: a memory for storing executable program code; and a processor for calling and running the executable program code from the memory, so that the electronic device implements a multimodal data management method as described in claim 9.

11. A computer program product or computer program, characterized in that, The computer program product or computer program includes computer instructions that, when executed by a processor, implement a multimodal data management method as described in claim 9.