Multi-modal interaction system based on large model driving

Through technical means such as the hexagonal cellular coding matrix and counterfactual correction module, the problems of insufficient cross-modal semantic alignment and resource allocation bottlenecks in the multimodal interaction system are solved, efficient and accurate multimodal interaction is achieved, the robustness and security of the system are improved, and flexible deployment in vertical fields is supported.

CN120508993AActive Publication Date: 2025-08-19BEIJING HONGYANGXUNTENG SCI TECH DEV CO LTD

Patent Information

Application Number
CN202510991452.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-19
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

The existing multimodal interaction system faces key defects such as insufficient cross-modal semantic alignment, lack of dynamic interaction and iterative reasoning capabilities, bottlenecks in resource allocation and computing efficiency, domain migration and security control defects, etc. when achieving efficient and precise interaction.

Method used

The hexagonal cellular coding matrix is used to store spatiotemporal features in layers, and a reconstructible alignment module is designed to generate adaptive weights through multi-head dot product attention and K-mean clustering to eliminate semantic gap errors; a counterfactual correction module is deployed, and the hidden variable relationship is stored using a structured causal graph database, and the adversarial sample optimization decision path is injected through a differentiable intervention engine; an elastic resource scheduler is integrated, and a parallel strategy based on the hardware-aware monitor is switched dynamically, and high-delay operator partition offload is achieved in combination with operator dependency analysis; a hot-swap field adaptation interface is built, and plug-and-play knowledge injection is supported through the encryption authentication unit and bandwidth isolation channel; a verification decision proof layer can be embedded to generate topological evidence graph linkage security filters.

Benefits of technology

Achieve cross-modal precise alignment, dynamic decision optimization and resource utilization improvement, breaking through the limitations of single visual understanding, achieving multimodal interaction efficiency at the human cognitive level, enhancing the robustness and security of the model, and supporting dynamic deployment and expansion of vertical fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508993A_ABST
    Figure CN120508993A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal interaction system based on large model driving, and relates to the field of multi-modal interaction. The system comprises an interaction agent terminal configured to receive multi-modal original data input by a user and output a fusion feedback result; the multi-modal gateway is in communication connection with the interaction agent terminal and is used for converting the fusion feedback result into a low-dimensional feature vector and eliminating a semantic gap between modals of the low-dimensional feature vector through a pre-trained attention mechanism; the large model driving engine is coupled with the multi-mode gateway and comprises a dynamically loaded multi-mode large model base supporting joint reasoning and a domain knowledge plug-in interface used for mounting a vertical domain fine tuning adapter; and the intention decision center is configured to analyze the output of the large model driving engine and generate an executable instruction. Unification of high precision, strong robustness and expandability of the industrial-grade multi-mode interaction system can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal interaction, and in particular to a multimodal interaction system driven by a large model. Background Art

[0002] The core goal of multimodal interaction systems is to enable natural and efficient information exchange between humans and machines. Their technological evolution is primarily driven by the following three dimensions: 1. The challenge of multimodal data complexity: Real-world interaction data encompasses multiple modalities, including text, speech, images, video, and point clouds, each with significant differences in data structure, feature space, and temporal resolution. Dynamic scenarios also require simultaneous analysis of both spatial structure and temporal evolution. 2. The evolution and limitations of large-scale model architectures: To address these challenges, hybrid expert architectures have become a hot topic. The SenseNova V6 model, with its 600 billion parameter MoE design, achieved breakthroughs in 64K token long-chain reasoning tasks, but computational cost remains a major bottleneck for industrial deployment. Pre-training on large-scale interleaved multimodal data has been shown to stimulate emergent capabilities. However, such data must cover spatiotemporal dynamic signals and cross-modal interaction features. Currently, the high cost of constructing such data and its variable quality limit model performance. 3. Refined requirements for application scenarios: Vertical fields such as medical and industrial require the system to have domain knowledge plug-in capabilities, and mobile devices and embedded systems require lightweight edge processing capabilities.

[0003] Given this background, current multimodal interaction systems face key drawbacks when achieving efficient and accurate interaction, including insufficient cross-modal semantic alignment, a lack of dynamic interaction and iterative reasoning capabilities, bottlenecks in resource allocation and computational efficiency, deficiencies in domain transferability and security controls, and limitations in evaluation systems. Therefore, to address these challenges, this paper proposes a multimodal interaction system driven by a large model. Summary of the Invention

[0004] Technical Purpose In order to solve the above problems, the purpose of the present invention is to provide a multimodal interaction system driven by a large model, aiming to eliminate the semantic bias caused by modal heterogeneity, break through the limitations of single visual understanding, optimize hardware resource utilization, and achieve cross-modal precise alignment, dynamic decision optimization and resource utilization improvement by constructing structured feature coding, flexible scheduling mechanism and counterfactual correction framework, and ultimately achieve multimodal interaction efficiency at the human cognitive level.

[0005] Technical Solution To achieve the above objectives, the present invention provides a large-model-driven multimodal interaction system. This solution uses a hexagonal cellular coding matrix to hierarchically store spatiotemporal features, designs a reconfigurable alignment module, and generates adaptive weights through multi-head dot-product attention and K-means clustering to eliminate semantic gap errors. A counterfactual correction module is deployed, which uses a structured causal graph database to store latent variable relationships and injects adversarial samples through a differentiable intervention engine to optimize decision paths. An elastic resource scheduler is integrated, which dynamically switches parallel strategies based on hardware-aware monitors and implements high-latency operator partitioning and offloading through operator dependency analysis. A hot-swappable domain adaptation interface is constructed, supporting plug-and-play knowledge injection through encrypted authentication units and bandwidth-isolated channels. A verifiable decision proof layer is embedded to generate a topological evidence graph and link security filters. This system achieves end-to-end efficient and accurate interaction through hardware-level coupling between a multimodal gateway, a large-model-driven engine, and an intent-based decision center.

[0006] In a first aspect, the present invention provides a multimodal interaction system driven by a large model, comprising: The interactive agent terminal is configured to receive multimodal raw data input by the user and output a fusion feedback result; A multimodal gateway, communicatively connected to the interactive agent terminal, includes: A modality-specific encoder group, configured to convert the fusion feedback result into a low-dimensional feature vector; A cross-modal alignment module that eliminates the inter-modal semantic gaps of the low-dimensional feature vectors through a pre-trained attention mechanism; A large model driving engine, coupled with the multimodal gateway, includes: Dynamically loaded multimodal large model base, supporting joint reasoning of text, voice, image, and video; Domain knowledge plug-in interface, used to mount vertical domain fine-tuning adapters; an intention decision center configured to parse the output of the large model driving engine and generate executable instructions, wherein the executable instructions are returned to the interactive agent terminal to drive the terminal device to perform a multimodal feedback operation; A resource scheduler is used to allocate computing resources to active task submodules in the large model driving engine in real time.

[0007] Furthermore, the interactive agent terminal integrates: Multi-core heterogeneous processors for independently running lightweight edge models; Protocol conversion bridge, compatible with Bluetooth / WiFi-6 / 5G communication protocol stacks.

[0008] Furthermore, the cross-modal alignment module further includes: A pre-trained feature aligner that optimizes the modality embedding space through contrastive learning; A reconfigurable cache pool dynamically stores high-frequency alignment parameters to reduce real-time computation load.

[0009] Furthermore, the domain knowledge plug-in interface adopts a hot-swappable physical connection structure, verifies the digital signature of the external adapter through an encryption authentication unit, and ensures the data exchange priority between the main model and the plug-in through a bandwidth isolation channel.

[0010] Furthermore, the intention decision center is deployed with an explainable proof layer for generating a topological evidence graph of the decision chain and a security compliance filter for blocking high-risk instructions based on a preset rule base.

[0011] Furthermore, it also includes a counterfactual correction module, which is connected downstream of the intention decision center and configured to generate adversarial samples to test the robustness of decisions and inject causal intervention signals to correct model output deviations.

[0012] Furthermore, the counterfactual correction module stores latent variable relationships in multimodal scenarios through a structured causal graph database, and optimizes the decision path through gradient backpropagation based on a differentiable intervention engine.

[0013] Furthermore, it also includes a dynamic computing offloading engine, which is configured to analyze the operator dependencies of large model reasoning and offload high-latency operator partitions to edge computing nodes.

[0014] Furthermore, the resource scheduler integrates: Hardware-aware monitor for real-time acquisition of GPU / FPGA power consumption and memory usage; An elastic computing allocator is used to dynamically switch between model parallelism and data parallelism strategies based on task latency requirements.

[0015] Furthermore, the multimodal gateway and the large model driving engine transmit feature vectors via a high-speed serial bus, and prioritize the processing of real-time modal streams via a hardware-level interrupt controller.

[0016] Furthermore, the dynamic computation offloading engine models computation offloading as a partially observable Markov decision process based on the task-resource bimodal hidden Markov model, and solves the optimal offloading strategy through variational Bayesian expectation maximization iteration, where the reward function is designed as:

[0017] Where, Output of the reward function; 、 and is the dynamic weight; is the delay sensitivity factor; is the total end-to-end delay of the task; is the total energy consumption of task execution; This is the benchmark value of full power of the mobile device; Accumulate jitter for the path; The maximum path jitter threshold allowed by the system.

[0018] Based on a task-resource bimodal hidden Markov model and a variational Bayesian expectation maximization framework, computation offloading is modeled as a multi-objective optimization problem. Through a unique path-accumulated jitter penalty term and dynamic sensitivity factor, latency, energy consumption, and network robustness are simultaneously optimized. This engine enables flexible and coordinated scheduling of edge-cloud resources, significantly improving the success rate of high-real-time tasks and achieving breakthroughs in mobile energy consumption control and network jitter adaptability.

[0019] Furthermore, a honeycomb topology projection module is included, which is used to decompose the spatiotemporal features into radial frequency components and tangential phase flows through a hexagonal unit spiral nested structure, and output a fused feature tensor using a Legendre polynomial orthogonal basis. The calculation formula of the fused feature tensor is:

[0020] Where, is the fusion feature tensor; is the hierarchical index of the cellular topology; is the high frequency retention coefficient of radial features; for Order Legendre polynomials; For the Radial frequency characteristics of layer honeycomb units; For the Layer tangential motion characteristics.

[0021] Through hexagonal honeycomb topology projection and Legendre orthogonal basis decomposition, the geometric distortion and information loss issues of multimodal features in the spatiotemporal dimensions are addressed. Its layered spiral structure decouples high-frequency texture and motion trajectories into radial-tangential orthogonal components, enabling continuous manifold embedding of cross-modal features. This matrix significantly improves the spatiotemporal modeling accuracy of dynamic scenes, effectively suppresses discontinuities at feature splicing boundaries, and reduces hardware resource consumption, laying a high-fidelity data foundation for subsequent cross-modal alignment.

[0022] In a second aspect, the present invention further provides a multimodal interaction method driven by a large model, the method being based on the system described in the first aspect, comprising: Receive the original multimodal data input by the user and generate a fused feature tensor through cellular topology projection; Based on the task real-time level and node resource status, the optimal offloading path is solved iteratively through variational Bayesian expectation maximization. Input the fused feature tensor into the large model base to generate an initial decision, retrieve the latent variable relationship through the structured causal graph database, inject adversarial samples using the differentiable intervention engine, and optimize the output instructions through gradient backpropagation; Computing tasks are partitioned and executed according to the offloading strategy. GPU / FPGA resource occupancy is dynamically collected through the hardware-aware monitor, and the elastic computing allocator switches between model parallelism and data parallelism.

[0023] In a third aspect, the present invention further provides a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the multimodal interaction method driven by a large model is executed.

[0024] The present invention eliminates the semantic gap of modal heterogeneity through a cross-modal dynamic alignment architecture, and uses a hexagonal cellular coding matrix to achieve high-fidelity fusion of spatiotemporal features; combines iterative reasoning with resource coordination mechanisms, relies on counterfactual correction modules to optimize decision paths, and uses elastic resource schedulers to achieve dynamic balancing of computing loads; further, through modular expansion and security control systems, with the help of hot-swappable domain adaptation interfaces, it supports agile deployment in vertical fields, while embedding a verifiable decision proof layer to ensure output compliance. This solution breaks through the modeling limitations of static alignment mechanisms for dynamic scenes, achieves geometric consistency expression of cross-modal semantics, overcomes the lack of iterative reasoning caused by single visual understanding, establishes closed-loop decision optimization driven by causal intervention, subverts the rigid resource allocation model, and constructs a hardware-aware adaptive scheduling framework, ultimately achieving the unity of high precision, strong robustness, and scalability of industrial-grade multimodal interaction systems.

[0025] Beneficial effects By implementing the multimodal interactive system based on large model drive provided by the present invention, the following technical effects are achieved: (1) This application provides the system with dynamic error correction capabilities by building a structured causal graph database and a differentiable intervention engine. By utilizing adversarial sample generation and gradient backpropagation, causal intervention in large model output decisions is achieved. This fundamentally reduces the decision-making risk of specialized tasks, enhances the model's robustness to distribution shifts and adversarial attacks, and generates a verifiable chain of evidence for decisions to meet security compliance requirements.

[0026] (2) A hardware-level design with encrypted authentication units and bandwidth isolation channels enables plug-and-play of vertical domain knowledge components. Physical isolation ensures the stability of the main model and avoids catastrophic forgetting caused by full parameter fine-tuning. This significantly shortens the deployment cycle of domain-specific modules, supports dynamic knowledge injection and cross-domain migration, and ensures the security and compatibility of the core system during expansion.

[0027] (3) Through hexagonal honeycomb topology projection and Legendre orthogonal basis decomposition, the geometric distortion and information loss problems of multimodal features in the spatiotemporal dimension are solved. Its hierarchical spiral structure decouples high-frequency textures and motion trajectories into radial-tangential orthogonal components, achieving continuous manifold embedding of cross-modal features. This matrix significantly improves the spatiotemporal modeling accuracy of dynamic scenes, effectively suppresses the discontinuity of feature splicing boundaries, and reduces hardware resource consumption, laying a high-fidelity data foundation for subsequent cross-modal alignment.

[0028] (4) Based on a task-resource bimodal hidden Markov model and a variational Bayesian expectation maximization framework, computational offloading is modeled as a multi-objective optimization problem. Through an innovative path-accumulated jitter penalty term and a dynamic sensitivity factor, latency, energy consumption, and network robustness are simultaneously optimized. This engine enables flexible collaborative scheduling of edge-cloud resources, significantly improving the success rate of high-real-time tasks and achieving breakthrough progress in mobile energy consumption control and network jitter adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to make the above-mentioned multimodal interaction system driven by a large model of the present invention more obvious and easy to understand, the following is a brief introduction to the drawings required for use in the specific implementation of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0030] Figure 1 Represents the system architecture diagram of this application; Figure 2 The flowchart of the present application method is shown. DETAILED DESCRIPTION

[0031] Example 1: A multimodal interactive system driven by a large model is provided. The system architecture is as follows Figure 1As shown, the system includes: an interactive agent terminal, deployed on user-side hardware devices, integrating a multi-core heterogeneous processor and a multi-protocol communication module. It receives multimodal raw inputs such as voice, images, and gestures and outputs fused feedback results. A multimodal gateway, connected to the agent terminal via a high-speed serial bus, includes a modality-specific encoder group and a cross-modal alignment module. The former converts each modal data into a low-dimensional feature vector, while the latter eliminates semantic gaps between modalities through a pre-trained attention mechanism. A large model driver engine establishes a hardware-level interrupt control channel with the gateway, and includes a dynamically loaded multimodal large model base and a domain knowledge plug-in interface. The latter adopts a hot-swappable physical structure and integrates an encryption and authentication unit. The intent decision center parses the large model output to generate executable instructions. It has a built-in explainability proof layer to generate a decision chain topology map and works with security and compliance filters to block high-risk instructions. The resource scheduler monitors GPU / FPGA power consumption and memory usage in real time and dynamically switches between model parallelism and data parallelism through an elastic computing allocator. Details are as follows.

[0032] The multimodal gateway workflow includes: Phase 1: Modal Coding: A Mel filter bank is used to extract spectral features from speech data, a convolutional neural network is used to extract spatial features from image data, and a 3D convolution kernel is used to separate spatiotemporal features from video data. The dimensionality of each modal feature vector is compressed to less than 5% of the original data volume.

[0033] Phase 2, cross-modal alignment: Optimize the embedding space using a pre-trained feature aligner, specifically: Compute the inter-modality similarity matrix based on contrastive learning; Dynamically store high-frequency alignment parameters through a reconfigurable buffer pool; Output a unified feature representation after eliminating semantic gaps.

[0034] The large model-driven engine operating mechanism includes: Dynamic loading mechanism switches the model base in real time according to the task type. The loading process triggers hardware-level interrupts to prioritize real-time modal flow.

[0035] The domain knowledge plug-in interface includes: physical structure, using PCIe hot-swappable interface to support plug-and-play of external adapters; security mechanism, encryption authentication unit verifies the digital signature of the adapter, bandwidth isolation channel ensures the data exchange priority of the main model; fine-tuning process, vertical domain knowledge injection only updates the adapter parameters to avoid retraining of the entire model.

[0036] The decision chain generation process includes: Step 1: The large model outputs preliminary intent labels; Step 2: The explainability proof layer constructs a decision evidence graph; Step 3: The security compliance filter matches the preset rule base.

[0037] The counterfactual correction process includes: Step 1: Retrieve the latent variable relationship in the structured causal graph database; Step 2: Differentiable intervention engine injects adversarial samples; Step 3: Adjust the final instruction based on the corrected output.

[0038] The hardware-aware monitor collects metrics including GPU memory occupancy, FPGA BRAM utilization, and DSP computing core load. The trigger condition is to initiate policy switching when the memory occupancy is greater than 85%.

[0039] A multimodal interaction method based on large model driving is provided. The method is based on the aforementioned system. The process is as follows: Figure 2 As shown, it includes: receiving the original multimodal data input by the user and generating a fused feature tensor through cellular topology projection; based on the task real-time level and node resource status, iteratively solving the optimal offloading path through variational Bayesian expectation maximization; The fused feature tensor is input into the large model base to generate an initial decision. The latent variable relationship is retrieved through a structured causal graph database, and adversarial samples are injected using a differentiable intervention engine. The output instructions are optimized through gradient backpropagation. The computing tasks are partitioned and executed according to the offloading strategy. The GPU / FPGA resource occupancy rate is dynamically collected through a hardware-aware monitor, and the model parallelism and data parallelism strategies are switched by an elastic computing allocator.

[0040] Example 2: Building on the previous examples, we modeled computational offloading as a partially observable Markov decision process based on a task-resource bimodal hidden Markov model. We iteratively solved the optimal offloading strategy using variational Bayesian expectation maximization, simultaneously optimizing latency, energy consumption, and network jitter sensitivity.

[0041] Defining hidden states

[0042] Where, It is a hidden state; Compute features for the task; Indicates the node resource status.

[0043] The reward function is designed as:

[0044] Where, Output of the reward function; 、 and is the dynamic weight; is the delay sensitivity factor; is the total end-to-end delay of the task; is the total energy consumption of task execution; This is the benchmark value of full power of the mobile device; Accumulate jitter for the path; The maximum path jitter threshold allowed by the system.

[0045] Perform a variational Bayes update:

[0046] Where, For the The hidden state of the iteration posterior distribution; For a given state Next action The observation probability of ,injecting the network topology prior; is the state transition probability; For Calculation of the expectation of a distribution.

[0047] For example, a mobile medical image analysis system processes lung CT sequences in real time.

[0048] Task A is lung nodule detection, which has high real-time performance and level is 3, the operator dependency graph is preprocessing → segmentation → detection, the computational cost 18.5 TFLOPS Task B is texture feature extraction, which has ordinary real-time performance and level is 1, the operator dependency graph is preprocessing → feature calculation, the amount of calculation 7.2TFLOPS; The node resource status at the moment is shown in Table 1.

[0049] Table 1. Summary of node resource status

[0050] The decision-making process simulation includes: Step 1, generating candidate offloading paths, as shown in Table 2; Table 2. Summary of candidate uninstall paths

[0051] Step 2: Calculate the reward function, as shown in Table 3; Table 3. Summary of reward function calculation

[0052] The decision output is to select the edge 2 path with the highest reward value.

[0053] The aggregated results of simulating 50 task executions are shown in Table 4.

[0054] Table 4. Summary of polymerization effects

[0055] According to the experimental table, under the constraint of the jitter threshold, the probability of exceeding the threshold drops to 5.2%; when the battery power is lower than 40%, the weight of the energy consumption item is automatically increased to 0.5; the edge 2 path wins due to low jitter and low energy consumption. Experimental data shows that the engine breaks through the Pareto frontier of multi-objective optimization. In a mixed load scenario of latency-sensitive, energy-sensitive, and network jitter-sensitive tasks, the system autonomously searches for the optimal strategy, and the trade-off boundary of latency, energy consumption, and robustness is systematically expanded compared to traditional methods. In the face of node resource fluctuations and task feature drift, the adjustment lag time of the offloading strategy is much lower than the industry benchmark, and the jitter suppression effect of the decision path shows strong stability. The refined modeling of the operator-dependent topology significantly reduces data waiting overhead, and the difference in resource utilization between the edge and the cloud converges to the ideal threshold, verifying the universal advantages of elastic scheduling.

[0056] Example 3: On the basis of the above-mentioned embodiments, a cellular topology projection mechanism is added to address the problem of semantic loss caused by the spatiotemporal separation of multimodal data. Through the spiral nested structure of hexagonal units, the spatiotemporal features are decomposed into radial frequency components and tangential phase flows. The Legendre polynomial orthogonal basis is used to achieve feature decoupling and reorganization, eliminating the boundary discontinuity error in the traditional grid structure.

[0057] The cellular encoding core is deployed in the FPGA chip, which contains 6 groups of DSP arrays, each group processing one hexagonal direction.

[0058] Per array integration: Radial frequency extractor, applying a modified Gabor filter:

[0059] Where, is the coordinate value in the rotating coordinate system; and is the adaptive bandwidth; is the fundamental frequency.

[0060] Tangential phase flow register, storing motion vector:

[0061]

[0062] Where, is the tangential motion vector; For the Structural confidence weight of the direction; For the hexagonal unit The image gradient vector in each direction; is the weight adjustment factor, preferably 0.7, which is used to control the sensitivity of similarity to weight; For the Direction sub-area With the central area The structural similarity index.

[0063] Reorganize the logic to generate fused feature tensors:

[0064] Where, is the fusion feature tensor; is the hierarchical index of the cellular topology; is the high frequency retention coefficient of radial features; for Order Legendre polynomials; For the Radial frequency characteristics of layer honeycomb units; For the Layer tangential motion characteristics.

[0065] Verification shows that this encoding structure significantly improves the geometric consistency of spatiotemporal features while achieving an average error similar to that of the above-mentioned embodiment. In tasks involving motion trajectory analysis, the system demonstrates the ability to accurately capture the laws of spatiotemporal evolution, effectively eliminating the telecentric distortion caused by traditional rectangular grids, and achieving industrial-grade accuracy for feature boundary continuity. The semantic loss rate of the multimodal fusion process decreases by orders of magnitude, especially in the decoupling and reorganization of high-frequency texture and low-frequency motion features, achieving distortion-free embedding and providing high-integrity input for the downstream alignment module. The encoder's parallel computing flow is highly consistent with the hexagonal topology, significantly reducing logic unit reuse conflicts, and the resource consumption curve meets superlinear compression expectations.

[0066] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product embodied on one or more computer-usable non-transitory storage media containing computer-usable program code.

[0067] The present invention can provide computer program instructions to a management platform of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the management platform of the computer or other programmable data processing device produce a device for implementing the system.

[0068] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device that implements the functions of the system.

[0069] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions of the described system.

Claims

1. A multimodal interactive system driven by a large model, characterized by: include: The interactive agent terminal is configured to receive multimodal raw data input by the user and output a fusion feedback result; a multimodal gateway, communicatively connected to the interactive agent terminal, configured to convert the fusion feedback result into a low-dimensional feature vector and eliminate inter-modal semantic gaps in the low-dimensional feature vector through a pre-trained attention mechanism; A large model driving engine, coupled with the multimodal gateway, including a dynamically loaded multimodal large model base supporting joint reasoning and a domain knowledge plug-in interface for mounting vertical domain fine-tuning adapters; The intention decision center is configured to parse the output of the large model driving engine and generate executable instructions, which are returned to the interactive agent terminal to drive the terminal device to perform multimodal feedback operations.

2. The system according to claim 1, wherein: The domain knowledge plug-in interface adopts a hot-swappable physical connection structure, verifies the digital signature of the external adapter through an encryption authentication unit, and ensures the data exchange priority between the main model and the plug-in through a bandwidth isolation channel.

3. The system according to claim 1, wherein: The intention decision center is deployed with an explainable proof layer for generating a topological evidence graph of the decision chain and a security compliance filter for blocking high-risk instructions based on a preset rule base.

4. The system according to claim 1, wherein: It also includes a counterfactual correction module, which is connected downstream of the intention decision center and configured to generate adversarial samples to test the robustness of decisions and inject causal intervention signals to correct model output deviations.

5. The system according to claim 4, characterized in that: The counterfactual correction module stores the latent variable relationships in multimodal scenarios through a structured causal graph database, and optimizes the decision path through gradient backpropagation based on a differentiable intervention engine.

6. The system according to claim 1, wherein: The multimodal gateway and the large model driving engine transmit feature vectors via a high-speed serial bus, and prioritize the processing of real-time modal streams via a hardware-level interrupt controller.

7. The system according to claim 1, wherein: It also includes a dynamic computation offloading engine that is configured to analyze operator dependencies for large model inference and offload high-latency operator partitions to edge computing nodes.

8. The system according to claim 7, characterized in that: The dynamic computation offloading engine models computation offloading as a partially observable Markov decision process based on the task-resource bimodal hidden Markov model, and solves the optimal offloading strategy through variational Bayesian expectation maximization iteration, where the reward function is designed as: Where, Output of the reward function; 、 and is the dynamic weight; is the delay sensitivity factor; is the total end-to-end delay of the task; is the total energy consumption of task execution; This is the benchmark value of full power of the mobile device; Accumulate jitter for the path; The maximum path jitter threshold allowed by the system.

9. The system according to claim 1, wherein: It also includes a honeycomb topology projection module for decomposing the spatiotemporal features into radial frequency components and tangential phase flows through a hexagonal unit spiral nested structure, and outputting a fused feature tensor using a Legendre polynomial orthogonal basis. The calculation formula of the fused feature tensor is: Where, is the fusion feature tensor; is the hierarchical index of the cellular topology; is the high frequency retention coefficient of radial features; for Order Legendre polynomials; For the Radial frequency characteristics of layer honeycomb units; For the Layer tangential motion characteristics.

10. A multimodal interaction method driven by a large model, characterized by: The method is implemented based on the system according to any one of claims 1 to 9: The method comprises: Receive the original multimodal data input by the user and generate a fused feature tensor through cellular topology projection; Based on the task real-time level and node resource status, the optimal offloading path is solved iteratively through variational Bayesian expectation maximization. Input the fused feature tensor into the large model base to generate an initial decision, retrieve the latent variable relationship through the structured causal graph database, inject adversarial samples using the differentiable intervention engine, and optimize the output instructions through gradient backpropagation; Computing tasks are partitioned and executed according to the offloading strategy. GPU / FPGA resource occupancy is dynamically collected through the hardware-aware monitor, and the elastic computing allocator switches between model parallelism and data parallelism.

Citation Information

Patent Citations

  • Large model interaction method and system based on multi-model collaborative dialogue

    CN119557842A

  • Method and system for collecting decision data by multi-modal large model driven intelligent agent

    CN119808006A

  • High-interpretability cross-modal extensible artificial intelligence system

    CN119990343A

  • Network transaction credit risk intelligent evaluation system and method based on multi-modal data fusion and deep learning

    CN120219047A

  • Intelligent decision support method for maritime safety information based on multimodal fusion network

    JP7623060B1

Cited By

  • Humanoid robot real-time interaction control system fusing monocular vision, voice and myoelectricity

    CN120886278A

  • Container mirror image pre-distribution method and system based on semantic layering and collaborative scheduling

    CN121309674A

  • A Container Image Pre-distribution Method and System Based on Semantic Layering and Cooperative Scheduling

    CN121309674B

  • Data decision generation system, method and device and electronic equipment

    CN121684026A

  • Video reasoning training data generation method based on key frame causal chain extraction

    CN121811185A