Multi-model dynamic image analysis system
Through heterogeneous multi-model collaborative inference and cross-modal spatial alignment technology, the problem of efficiency and accuracy imbalance in dynamic image analysis is solved, efficient and reliable processing of multimodal data is achieved, and hardware resource utilization and reliability of key scenarios are improved.
Patent Information
- Application Number
- CN202510636940.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-18
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art has problems such as single model limitations in dynamic image analysis, resulting in imbalance of efficiency and accuracy, lack of multimodal data collaboration mechanism, low hardware resource utilization and insufficient reliability in key scenarios, especially in medical image analysis and traffic monitoring, it is difficult to achieve efficient and reliable multimodal data fusion and diagnosis.
Using heterogeneous multi-model collaborative inference, cross-modal spatial alignment and dynamic weight allocation, cross-modal KV cache optimization and knowledge base constraint technology, high-precision and low-latency multi-scene dynamic analysis is achieved through the heterogeneous fusion of multi-modal large language model and object detection model.
It realizes efficient collaborative processing of multimodal data, dynamic weight allocation and adversarial cache optimization, improves the efficiency, accuracy and reliability of the system, reduces manual review dependence, improves hardware resource utilization, and enhances the reliability of key scenarios.
Smart Images

Figure CN120451753A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence vision technology and involves methods for collaborative computing of large multimodal models, dynamic image analysis, and the construction of intelligent decision-making systems. It is particularly applicable to scenarios such as medical imaging diagnosis, real-time traffic monitoring, and dynamic clinical decision support. Specifically, it involves heterogeneous computing resource optimization, cross-model cache sharing, dynamic fusion of multimodal features, and decision verification technology based on knowledge graphs. Background Art
[0002] Existing technologies have the following core pain points in dynamic image analysis: 1. The limitations of a single model lead to an imbalance between efficiency and accuracy.
[0003] Traditional systems rely on single visual models (such as Qwen-VL or YOLO series) to process medical images or traffic monitoring data, making it difficult to balance processing speed and diagnostic accuracy. For example, medical image analysis requires multiple manual reviews to ensure reliability, leading to diagnostic delays.
[0004] 2. Lack of multimodal data collaboration mechanism and cross-modal alignment capabilities.
[0005] Existing systems lack the ability to integrate voice, physiological data, and images in real time: (1) During medical consultations, physiological data such as patient voice, emotion, heart rate, and blood oxygen are not dynamically associated with imaging features, resulting in the diagnostic process relying on manual integration; (2) Multimodal models (such as Qwen-VL) often suffer from the problem of “inattention”, such as lack of contextual coherence in the description of multiple related images, or incorrect identification of key lesions (such as misclassifying “crabapple flower” as other species).
[0006] 3. Low hardware resource utilization Traditional architectures do not fully utilize the heterogeneous computing capabilities of the CPU and do not implement multi-model parallelism or cache management.
[0007] 4. Insufficient reliability in key scenarios The existing system is vulnerable in the following scenarios: (1) Handling controversial medical diagnoses: There is a lack of a secondary verification mechanism for analytical results with confidence levels below the threshold, which can easily lead to the risk of misdiagnosis; (2) Dynamic medical response: Failure to trigger emergency procedures based on the patient's real-time physiological data (e.g., heart rate > 100 or blood oxygen < 93), resulting in delayed treatment of critical situations; (3) Lack of model co-verification: The consistency of attention regions of different visual models (such as Qwen-VL and YOLOv11) is not constrained, resulting in lesion localization errors.
[0008] Existing technical solutions have the following drawbacks: (1) Agent frameworks (such as Coze and Dify) are difficult to flexibly adapt to custom multi-model pipelines; (2) Simply relying on multiple models in series (such as VL model + QWQ reasoning) does not solve the problems of feature alignment, cache sharing and dynamic weight allocation, resulting in performance degradation in complex scenarios.
[0009] In summary, there is an urgent need for a system architecture that combines heterogeneous multi-model collaboration, dynamic resource optimization, multimodal feature fusion, and an adaptive verification mechanism to break through the bottlenecks of existing multimodal technologies in efficiency, accuracy, and reliability. Summary of the Invention
[0010] The present invention proposes a multi-model dynamic image analysis system, which realizes high-precision and low-latency multi-scene dynamic analysis through heterogeneous multi-model collaborative reasoning, cross-modal spatial alignment and dynamic weight distribution, cross-model KV cache optimization and knowledge base constraints.
[0011] The system architecture is as attached Figure 1 Shown, including: Task decomposition and reconstruction (including but not limited to implementation by embedded small model SLM or direct user input), multimodal or target detection vision module (e.g., a multimodal large model LLM and a target detection model, or, for example, a forward analysis multimodal LLM module and a different target elimination LLM verification module), general precision LLM reasoning enhancement module, enhanced knowledge base, high-precision LLM reasoning and adjudication module; And the system performs: a) The task decomposition and reconstruction module decomposes or heterogeneously reconstructs the input multimodal instruction execution task, including but not limited to decomposing it into a forward detection task, a reverse target elimination verification task, and a fine-grained reasoning enhancement task through embedded model reasoning, or decomposing it into target area detection, target attribute analysis, and reasoning enhancement tasks.
[0012] b) The multimodal or target detection vision module performs heterogeneous processing and fusion, including but not limited to processing the forward detection task through a high-precision multimodal model and outputting structured analysis results, while the target detection module generates a verification heat map according to the heterogeneous elimination verification task, performs reverse target elimination and performs spatial consistency comparison with the forward results, and then performs fusion through an embedded low-precision multimodal model.
[0013] The multimodal LLM module retrieves dynamic knowledge index units when analyzing images, realizing retrieval-generation collaboration.
[0014] c) The reasoning enhancement module then performs enhancement work on the analysis results reports of each of the above parts based on task decomposition or instruction intent, including but not limited to logical verification, in-depth interpretation, reasoning, or targeted fine-grained feature enhancement. It generates preliminary diagnosis or event analysis conclusions through multimodal feature fusion and can optionally automatically trigger subsequent verification processes based on confidence thresholds. d) Optionally, the adjudication optimization module reviews controversial multimodal analysis results (e.g., those with a confidence level below 95% or that trigger sensitive vocabulary), generates enhanced conclusions, analyzes the system's operating status, and dynamically reconstructs the KV cache structure.
[0015] For example, the multimodal and target detection vision module may include at least one multimodal large model module and at least one target detection training module (such as YOLOv8). One way of heterogeneous fusion thereof may be: Step S1: The pre-processing module performs task decomposition and instruction intention attention area positioning; Step S2: The object detection module (e.g., YOLOv8) generates preliminary object detection results (e.g., searching for suspected holes in containers). Step S3: The multimodal LLM verification module ensures the reliability of the results through a heterogeneous verification mechanism (e.g., finding the hole-shaped loading and unloading port required for normal operation), and performs feature extraction and probability prediction on the detection task, as shown in the attached figure. Figure 2 As shown; Step S4: The reasoning module summarizes the S2 and S3 reports, performs semantic enhancement on the features, and deploys a large language model to perform logical verification and in-depth interpretation of the detection results.
[0016] Or, the reverse process is to first perform the forward analysis by the multimodal LLM, as shown in the attached figure. Figure 3 shown.
[0017] For example, another way of heterogeneous fusion can be: a) the multimodal large model module receives image input and generates a pseudo label of the target area in combination with the text guide in the knowledge base; b) the pseudo-label generation module performs three-dimensional spatial consistency verification to select reliable labels with a confidence level higher than a certain threshold; c) The target detection training module fuses strong labels and weak labels to optimize the target detection training module (such as YOLOv8) model.
[0018] Target detection modules (such as YOLOv11) and multimodal large models can also coordinate positioning: Fast region detection: Deploy the target detection module to perform high-precision target detection (such as lung nodules, tumors, and lesion areas) on images, and output the bounding box coordinates (Bbox), category probability, and key point location information of the lesion area; Feature region extraction (optional): Based on the detection results of the target detection module, the system automatically crops or enhances the image blocks of the target area (such as applying CLAHE histogram equalization to low-contrast areas) to generate high-resolution image sub-images; Attention guidance optimization: The detection results of the target detection module are used as the visual attention prior of the multimodal model (such as Qwen-VL-72B). The coordinate information of the lesion area is injected into the model through position encoding (such as Sine Positional Encoding) in the Transformer encoding layer, forcing it to focus on key locations. Joint reasoning process: After receiving the enhanced regional features and detection confidence, the multimodal model dynamically adjusts its contextual understanding weight (for example, in chest X-ray diagnosis, it assigns a double weight to the detected lung shadow area) to generate a more accurate diagnostic conclusion.
[0019] Another example of heterogeneous fusion is to achieve high-precision, low-latency container damage detection and integrity assessment by building a three-level collaborative architecture consisting of a "72B damage detection model, a 27B integrity analysis model, and a 671B inference hub." The system first utilizes the 72B multimodal model to perform fine-grained damage identification on containers. This model fuses visual images (surface texture captured by an RGB camera) with point cloud data (3D structural information from laser scans) and employs a multimodal feature fusion layer (a product of a vision_transformer and a pointnet) to locate 12 types of damage, including dents, perforations, and rust. Monte Carlo Dropout is then used to quantify uncertainty, marking only suspected areas with a confidence level ≥ 0.7.
[0020] Subsequently, the integrity analysis is performed using the lower-precision 27B multimodal model. The model performs 3D topology inference on the six faces, models the container as a grid unit, predicts the integrity probability of each unit through the Transformer layer, and uses morphological closing operations to connect continuous complete areas, ultimately outputting a completeness heat map with fuzzy boundaries. The detection results of the two-level model will be transmitted to the 671B reasoning center in the form of structured text, as shown in the attached file. Figure 5 The model uses a layer of evidence integration to compare conflicting information. For example, if the damage model indicates a perforation on the left side (confidence level 0.82) while the integrity model is rated B (missing 12% of the area), the built-in physical rule engine (e.g., "a perforation necessarily results in integrity ≤ C") is triggered for arbitration, ultimately generating a report that includes the damage location, repair recommendations, and transportation risk level.
[0021] For example, another way of heterogeneous fusion is to use the same model to implement a single-model two-stage forward and reverse detection process. For example, by reasoning twice in different modes on the same model, the coordination of container damage detection and integrity assessment can be achieved. Specifically, the system uses a large multimodal model with 72 billion parameters (such as Qwen-VL-72B), and the architecture design includes feature sharing and dynamic task switching. In the first forward reasoning stage, the model receives multimodal input of the container (including RGB images and lidar point cloud data), extracts common visual features (such as surface texture, three-dimensional structure, material reflection characteristics, etc.), and identifies the damaged area of the container.
[0022] After completing the forward detection, the model enters the reverse analysis phase. At this point, the system switches the task mode of the second half of the model through a dynamic switch module, converting it from "damage detection" to "integrity assessment." The reverse phase directly reuses the common features extracted by the shared layer in the forward reasoning, eliminating the need for repeated calculations and significantly reducing resource consumption. In this mode, the model actively masks the damaged areas that have been marked in the forward phase, and through the built-in three-dimensional spatial reasoning capabilities, it models the container as a gridded topological structure, predicting the integrity of the undamaged areas unit by unit. Specifically, the model focuses on the "unmarked" areas through the attention mechanism, combines the morphological closing operation algorithm to integrate continuous and complete areas, and performs logical constraints based on preset physical rules (such as "the presence of perforations will inevitably lead to a local integrity score lower than level C" and "five consecutive grid cells are intact and considered a complete surface", etc.), to ensure that the reverse analysis results and the forward conclusions form a complementary verification to avoid missed detection or misjudgment. The output of the reverse phase includes a heat map of the complete surface, an integrity level score for each surface (such as AF level), the coordinates of conflicting areas that require manual review (for example, when an area with a high confidence level in the forward direction is mistakenly judged as intact in the reverse direction), and a container transportation risk assessment based on comprehensive analysis (such as whether transportation is allowed, whether special reinforcement is required, etc.). Ultimately, the structured results of the two reasonings are integrated into a standardized report and connected to the reasoning model. This solution can achieve the dual, heterogeneous fusion functions of "damage identification" and "integrity inference" through a single-model architecture design, rather than relying on the collaboration of multiple independent models. In addition, the system improves the reliability and compliance of the output results through a built-in physical rule engine and conflict detection mechanism (for example, when an area with a forward confidence level exceeding 0.5 is automatically triggered to re-analyze when it is judged as intact in the reverse direction).
[0023] The system also includes a cross-model multimodal cache management system, which is characterized by including: Global cache coordinator, cross-model adaptation layer, cache prediction module and hierarchical storage unit; Monitor the KV cache status of each model in real time and implement cross-model KV cache sharing and dynamic compression. Establish a cross-model KV cache collaborative workflow between multimodal models, including the following steps: a) Extracting key-value (KV) caches of each attention layer from multiple heterogeneous multimodal models; b) Projecting the KV caches of features of different dimensions into a unified vector space through a cross-model adaptation layer; c) Analyze access patterns based on the cache prediction module; d) Implement a tiered storage strategy based on the prediction results (e.g., high-frequency hotspot cache resides in GPU memory, medium-frequency cache is stored in CPU-GPU shared memory, and low-frequency cache is compressed and stored in a solid-state drive array). The recommended steps for a KV cache collaborative workflow are as follows: Step 1: Cache preprocessing 1. Each modal model completes the forward calculation; 2. Extract the KV cache of each attention layer; 3. Convert to a unified format (e.g., dimension 2048) through an adaptation layer.
[0024] Step 2: Collaborative Storage 1. Write the processed KV cache to the shared memory pool; 2. Update the metadata of the cache mapping table; 3. Perform compression operation.
[0025] Step 3: Smart Scheduling 1. The monitoring module collects real-time information on GPU memory usage, model inference progress, and cache access hotspot distribution; 2. Dynamic decision engine based on preset strategies, such as: A. High-frequency cache: kept in GPU memory; B. IF cache: stored in CPU-GPU shared memory; C. Low-frequency cache: compressed and stored in SSD.
[0026] Step 4: Reuse across models When Model B needs to access Model A's cache: (1) Query the cache mapping table to obtain location information; (2) If the dimensions do not match, the adaptation layer conversion is triggered.
[0027] Preferably, the intelligent cache management system includes a cross-model KV cache sharing protocol, which maps the cache space of different models to a unified memory pool through hash indexing, such as allowing Qwen-VL-72B and Gemma-27B to share the intermediate results of the visual feature encoding layer.
[0028] Preferably, the intelligent cache management system includes a dynamic transfer learning module to retrain low-confidence cache data. Figure 4 As shown. Including but not limited to: (1) Low-confidence data screening unit: identifies cached data blocks that need to be retrained through a confidence threshold mechanism; (2) Incremental learning framework: Based on Domain-Adaptive Fine-Tuning technology, parameter fine-tuning is performed on the screened data; (3) Knowledge distillation verification: The model output after transfer learning is compared with the expert knowledge base, and the main cache is updated only when the F1-score ≥ 0.85 is met.
[0029] To address the issue of attention drift with multi-image input, the system employs a serialized processing mechanism, forcing the multimodal model to analyze each image individually. Structured metadata is then appended to each image to anchor the context. For example, given the frequent inattention issues with VL models, this system inserts loops to ensure the multimodal model processes only one image at a time. For example, a subheading is pre-set for each image description, such as "The following is a description of image {{ n}} / {{ total}}." Images are separated by Markdown delimiters.
[0030] Preferably, the deep parsing layer includes a model parallel deployment solution, which implements heterogeneous reasoning in the processor and memory (the activation expert layer of MOE can still be in the graphics card) through parameter sharding technology (e.g., dividing DeepSeek into 8 logical units), and realizes memory affinity optimization under a dual-core NUMA architecture.
[0031] Specific implementation cases The technical solutions adopted in the specific implementation case 1 are as follows: 1. System architecture and core modules Heterogeneous computing resource optimization architecture: The system is based on a 4029 ten-card platform, equipped with 8×2080Ti 22G GPUs + 1×3090 24G GPU + dual-NUMA CPU (this machine has two processors, forming a dual-NUMA), 1024GB of DDR4 memory, and realizes dynamic resource allocation and collaborative computing through the following modules: (1) The first part of the system is the preprocessing model O, which occupies one graphics card. This model is a preprocessing buffer layer (which can be a 7b multimodal visual small model and an embedded small model). It first performs quality filtering on the image and can also roughly locate the attention area. Then it disassembles and reconstructs the task, including but not limited to disassembling it into a forward detection task and a heterogeneous complementary classifier verification task. Then, the forward detection task is handed over to the high-precision multimodal model in the second part, and the heterogeneous complementary classifier verification task is handed over to the relatively low-precision second visual model in the third part. (2) The second part of the system is the high-precision multimodal model A: it has four graphics cards, numbered graphics cards 2 to 5, with a total of 88G video memory. It runs a multimodal vision model such as qwen-vl-72b at a very high speed and can process about 20 requests in parallel.
[0032] (3) The third part of the system is the second visual model or target detection model B with relatively low precision: 1 graphics card deployment (number 6), used to perform heterogeneous tasks, such as the previous visual model A outputs the probability map of the lesion area, and here the visual model B generates the Grad-CAM heat map, by comparing the heat maps of the two models. Figure 1 Consistency constrains the model to focus on the same medical markers (such as cancer cell nuclei). (4) The fourth part of the system is the fast but low-precision inference model C: Two graphics cards (numbered 7-8) run an analytical inference model, such as qwq32b, as a fast inference layer, connected to the preprocessing buffer layer for further analysis of the image features extracted by the visual model. The speed is roughly matched with the multimodal model, and can process about 40 requests per minute.
[0033] (5) The fifth part of the system is the cross-model kvcache sharing and compression system D: it implements the intelligent cache layer cache management system, kv cache cache compression, kv cache cache decompression functions, connects the hybrid inference layer and the deep model, and implements prediction-based cache management.
[0034] (6) Intelligent cache management system: A. Implement cross-model KV cache sharing and dynamically allocate video memory through hash indexing and LSTM prediction models; B. Implement retraining on low-confidence cache data (such as rare disease cases) and support video memory pressure prediction and automatic compression.
[0035] (7) System Part 6: A dual-core processor, a 3090 24G graphics card, and over 1024GB of memory are used to run the full-featured DeepSeek 671B Q2.51 dynamic quantization version through a heterogeneous computing solution to further enhance the interpretation of approximately 10% of controversial multimodal results or important judgment work triggered by sensitive content lexicons. However, this part is relatively slow, with a concurrent capacity of around 5 channels, and can only process two or three requests per minute.
[0036] (8) Part 7 of the system also has a gateway: the model gateway aipx-gateway uses the ports of each part of the system as its own backend, and then provides the client with a unified OpenAI API as the entrance to the cluster. (9) The main process of the gateway aipx-gateway is as follows: First, the client requests aipx-gateway, uploads pictures and questions to the gateway, and aipx-gateway needs to decode the pictures carried in the request, and use the pictures to request the multimodal model (such as qwen2.5vl) to obtain the description of the image, and generate image descriptions one by one (with subtitles and Markdown delimiters), and then send the analysis report of the picture together with the user's original request to the inference model to generate preliminary conclusions, and trigger the collaborative verification or deep analysis process according to the confidence level. As shown in the attached Figure 6 shown.
[0037] (10) Dynamic routing: A regular request: routed to the embedded model disassembly and reconstruction task → Qwen-VL-72B forward → YOLOV11 target detection → GEMMA3-27B out-of-phase → QWQ-32B preprocessing layer reasoning → DeepSeek depth analysis layer; B. Controversial results: trigger the collaborative verification layer and medical prior constraint module.
[0038] (11) Multi-model pipeline optimization A. Prompt word engineering: (a) Add subtitles and separators to multiple image descriptions (e.g., "Description of image 1 / 3:\nLung CT shows a nodule in the left lobe with blurred boundaries\n---"); (b) Additional prompt words: "Please note that the above description may not be accurate and needs to be verified in combination with the context and medical knowledge."
[0039] B. Result integration: Encapsulate the visual model output, inference conclusion, and confidence information into a structured response (such as JSON format) and return it to the client.
[0040] The aforementioned multiple graphics cards still have a considerable amount of free space, which can be used to deploy more multimodal models or target detection models, or for KV CACHE caching.
[0041] The above design makes full use of the performance of this machine.
[0042] Specific implementation case 2 (1) Feature 1 of this embodiment: The hardware allocation of each module in this embodiment is as follows: (1) Preprocessing buffer layer module, whose hardware configuration is graphics card 7 (22GB) + lightweight 7B model, can realize the functions of data standardization, quality filtering, and metadata tagging.
[0043] (2) Multimodal visual processing cluster module, whose hardware configuration is graphics card 1-4 (4×22GB) + Qwen-VL-72B, can realize the functions of medical image analysis and lesion area probability map output. (3) Collaborative verification layer module, whose hardware configuration is graphics card 8 (22GB) + Grad-CAM / YOLOv11, can realize the functions of heat map generation, consistency comparison, and medical prior constraints. (4) Low-precision fast reasoning layer module, whose hardware configuration is graphics card 5-6 (2×22GB) + QWQ-32B, can realize the functions of generating preliminary diagnostic conclusions and confidence assessment. (5) High-precision deep analysis layer module, whose hardware configuration is processor A, B + 1024GB memory + graphics card 9RTX3090 + DeepSeek 671B model, can realize the functions of controversial result review and knowledge graph enhanced decision-making. (6) Intelligent cache management system module, whose hardware configuration is processor A (512GB memory), can realize KVCache sharing, compression / decompression, and video memory optimization functions. (II) Feature 2 of the embodiment: Multimodal dynamic consultation process 2.1 Client Requests Access (1) The user uploads a composite request containing medical images, voice complaints, and physiological data to the unified gateway aipc-gateway through the medical terminal; (2) aipc-gateway parses the request, separates modal data such as images, voice, and text, and starts an encrypted transmission channel.
[0044] 2.2 Multimodal Feature Extraction and Fusion (1) Visual processing flow: A.aipc-gateway calls the Qwen-VL-72B on graphics cards 1-4 to perform multi-scale analysis on medical images and output image feature vectors and lesion area probability maps. B. Synchronously start the collaborative verification layer of graphics card 8 and generate the lesion area bounding box (Bbox) and Grad-CAM heat map through YOLOv11; C. Use the ICP algorithm to compare the consistency of the probability map and the heat map. If the positioning error is greater than the threshold (for example, the error of the cancer cell nucleus is greater than 2 pixels), the review mechanism is triggered.
[0045] (2) Voice and physiological data processing: A speech module extracts the emotional index (e.g., anxiety index > 85 is marked as a high-concern case); B Physiological sensor data (heart rate, blood oxygen) is accessed in real time, and abnormal values (such as heart rate > 100 or blood oxygen < 93) directly activate the deep analysis channel.
[0046] 2.3 Dynamic Reasoning and Decision Making (1) Preprocessing and feature compensation: A preprocessing layer (GPU 7) verifies the feature distribution (such as variance and maximum value) of the visual model output and performs compensation or filtering on low-confidence results; B uses a loop mechanism to generate descriptions for multiple related images one by one (e.g., “Description of the 1 / 3 image: Lung CT shows a nodule in the left lobe with blurred boundaries”).
[0047] (2) Fast inference layer processing: A.aipc-gateway integrates the standardized visual features, speech emotion labels, and physiological data and transmits them to the QWQ-32B of graphics card 5-6; B. QWQ generates preliminary diagnostic conclusions (e.g., “left ventricular wall motion abnormalities are 92% correlated with chest pain”) and evaluates the confidence distribution (e.g., Monte Carlo sampling).
[0048] (3) In-depth analysis of controversial scenarios: A. If the confidence level is less than 95% or a sensitive word library (such as "malignant tumor") is triggered, the request is routed to the DeepSeek model of processor B; B. DeepSeek combines knowledge graphs (over 100,000 medical relationship triplets) to generate differential diagnosis pathways (such as three possible causes and evidence support) and initiates multi-expert arbitration (≥3 sub-models unanimously approved).
[0049] 2.4 Decision Output and Interaction (1) Generate a structured diagnostic report, including: A. Lesion area location coordinates and confidence level; B. Multimodal feature fusion weight matrix (visual display); C. Dynamic decision tree (interactive topology diagram, with medical evidence sources marked).
[0050] (2) High-risk cases will be simultaneously sent to the doctor's workstation as early as possible, along with follow-up suggestions (e.g., "further confirmation of the patient's recent medication history is required").
[0051] (III) Feature 3 of the embodiment: Implementation of intelligent cache management system 3.1 Cross-model KV Cache Sharing (1) Hash index mapping: A. Processor A maps the intermediate results of Qwen-VL-72B and QWQ-32B (such as the output of the visual feature encoding layer) to a shared memory pool; B. Reduce cache conflicts and improve cross-model reuse rates through dynamic hashing algorithms.
[0052] (2) Compression and decompression process: A. Use the LSTM prediction model to predict future cache demand and compress low-priority cache data (compression ratio ≥ 5:1); B. When a request arrives, monitor the memory usage through a sliding time window. If the prediction is insufficient, decompress high-value cache items first (decompression delay ≤ 20ms).
[0053] 3.2 Dynamic Optimization of Video Memory (1) Gradient Accumulation Cache (GAC): Reduces peak memory usage by 40% by accumulating small batches of gradient updates. (2) Resource recovery during idle period: A. Start cache preloading (preheating common model parameters); B. Perform model self-diagnosis (such as memory leak detection and calculation accuracy verification).
[0054] (IV) Feature 4 of the embodiment: Adversarial Cache Optimization Design 4.1 CacheGAN training and reasoning, as shown in the attached Figure 7 shown.
[0055] (1) Generator: predicts future demand based on historical cache access patterns and generates candidate cache content; (2) Discriminator: evaluates the future value of candidate content (e.g., relevance to the current task, reuse probability); (3) Adversarial training: Optimize the generator by minimizing the discriminator loss function to improve prediction accuracy.
[0056] 4.2 Neural Cache Topology Management (1) Construct a semantic graph-based cache structure, where nodes represent cache items and edges represent associations (e.g., “pulmonary nodule image” and “QWQ reasoning conclusion”); (2) Use graph neural networks (GNN) to dynamically adjust node weights and prioritize high-association caches.
[0057] 4.3 Memory Replay Mechanism (1) Regularly replay edge cases (such as rare disease images or controversial diagnoses) and update cache weight coefficients through reinforcement learning to improve the model's adaptability to long-tail scenarios.
[0058] (V) Feature 5 of the embodiment: Dynamic decision tree generation and update (1) In the specific implementation case 2, this system is used to assist outpatient clinics. It has a multimodal dynamic decision tree generation engine. Based on the LLM reasoning results and the patient's multimodal data, it constructs a tree structure containing priority screening paths, differential diagnosis nodes and follow-up questioning suggestions, and constrains the medical rationality of decision branches through knowledge graphs.
[0059] (2) Dynamic medical consultation decision support process: A. Voice recognition of patient complaints, simultaneous transmission of physiological data (heart rate, blood oxygen) by wearable devices, and multimodal alignment achieved through embedded small models.
[0060] B. Receive composite requests containing images through the multimodal gateway, decode and extract image feature vectors, input multimodal features into the fast inference layer for joint analysis, and generate preliminary diagnosis conclusions and their confidence distributions; C. Start the collaborative validation process: compare the lesion probability of the multimodal model with the visual model and calculate the consistency index (needs to be ≥90% to pass); D. In dispute scenarios, the heterogeneously deployed DeepSeek model is combined with the knowledge graph for secondary reasoning, generating a final diagnosis report and treatment recommendations. E.LLM integrates emotional and physiological indicators to generate a personalized consultation path (for example, prioritizing heart disease screening). It also dynamically generates a list of follow-up questions ("Have you had chest pain recently?") and pushes them to the doctor's interface. LLM generates a dynamic diagnostic tree, reducing the doctor's cognitive load.
[0061] (3) Dynamic decision tree generation methods may include: A. Initial node creation: Generate a candidate diagnosis set based on the patient's chief complaint keywords (such as "chest pain") and abnormal physiological indicators (such as ST segment elevation >1mm); B. Branch weight assignment: Candidate diagnoses are ranked by urgency through knowledge graph semantic similarity calculation (e.g., Word2Vec embedding distance) to generate a tree structure containing priority labels; C. Real-time update mechanism: Dynamically expand branches based on newly entered follow-up questions or examination results (e.g., adding aortic dissection branches when chest pain is accompanied by hypotension); D. Visual output: Convert the decision tree into interactive interface elements for doctors (such as a topological map with confidence levels), and simultaneously annotate the medical evidence source for each node.
[0062] (4) Initial decision tree construction A. Keyword extraction: Extract core symptoms (e.g., "chest pain") and physiological abnormalities (e.g., "ST segment elevation > 1 mm") from the patient's chief complaint; B. Candidate diagnosis generation: Generate a set of candidate diagnoses (e.g., "myocardial infarction" and "aortic dissection") by calculating semantic similarity on the knowledge graph (e.g., Word2Vec embedding distance); C. Branch weight allocation: Based on the urgency sorting, a tree structure with priority labels is generated.
[0063] (5) Real-time updates and interactions A. New data access: Expand branches based on follow-up questions or examination results (e.g., adding a "pulmonary embolism" branch); B. Visualization output: Convert the decision tree into an interactive topological diagram, annotating the confidence level and medical evidence source of each node (e.g., citing clauses from the Clinical Diagnosis and Treatment Guidelines); C. Constraint verification: Ensure that branches comply with clinical specifications and avoid logical conflicts through knowledge graph path search.
[0064] Through the above-mentioned specific implementation methods, the present invention realizes efficient collaborative processing of multimodal data, dynamic weight allocation, adversarial cache optimization and medical knowledge constraints, and solves the core problems of existing technologies in terms of efficiency, accuracy and reliability. Beneficial effects
[0065] The multi-model dynamic image analysis system proposed in this paper, through core technologies such as multi-model collaborative reasoning, dynamic weight allocation, adversarial cache optimization, and knowledge constraints, has the following significant advantages over existing technologies: 1. Achieve a balance between efficiency and precision (1) Heterogeneous collaborative reasoning: Through the division of labor of heterogeneous models with non-unidirectional verification and heterogeneous fusion, concurrent requests can be quickly processed, the reasoning model generates explainable conclusions, and the verification model ensures the reliability of key scenarios, forming a layered processing architecture.
[0066] (2) Dynamic weight allocation mechanism: The fusion weight is adjusted in real time according to the modal confidence level, avoiding the subjectivity of manual weight setting and reducing the reliance on manual review.
[0067] 2. Significantly optimized resource utilization (1) Dynamic memory management: Through cross-model KV Cache sharing, gradient accumulation cache technology and intelligent memory allocation strategy, it reduces memory fragmentation and improves hardware resource utilization.
[0068] (2) Heterogeneous computing collaboration: It fully utilizes the heterogeneous architecture of dual-core CPUs and multiple GPUs to achieve parallel processing of the deep parsing layer and lightweight models, thus avoiding idle computing resources.
[0069] 3. Enhanced system stability and security (1) Automatic review of dispute scenarios: For results with abnormal confidence or triggering sensitive vocabulary, the high-precision deep analysis layer (such as the Full-precision DeepSeek model) is automatically activated, and multi-path diagnostic suggestions are generated in combination with the knowledge graph to reduce the risk of misdiagnosis.
[0070] (2) Dynamic cache optimization: Improve cache hit rate and reduce repeated calculations through adversarial cache design (such as CacheGAN and neural cache topology), while ensuring data privacy and transmission security.
[0071] 4. Improved scalability and compatibility (1) Unified service gateway: A gateway compatible with the OpenAI API format enables unified access and scheduling of multi-model services, supporting flexible expansion of new models or business scenarios.
[0072] (2) Modular design: Each functional module (such as the preprocessing layer, inference layer, and verification layer) can be independently optimized or replaced to adapt to different hardware configurations and business requirements.
[0073] 5. A multi-modal analysis model with deep thinking is implemented, and the test is as shown in the attached Figure 8 With attached Figure 9 As shown in the figure, when given a photo of flowers and asked to recommend related ancient poems, after the task of disassembly, reconstruction, enhancement and fusion, the recommended related ancient poems finally given by the large model are obviously closer to the content of the photo than the results directly given by the 72B multimodal model, showing a significant improvement effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] The system and method of the present invention are further illustrated by the following figures to illustrate their architecture and core technology implementation: Figure 1 : Architecture diagram 1) Feature Extraction Cluster: This cluster consists of a multimodal visual processing cluster (Qwen-VL-72B) and object detection and heterogeneous complementary layers, demonstrating the heterogeneous localization mechanism of the visual model and the dedicated detection model. 2) Decision-making and reasoning pipeline: This includes a low-precision fast inference layer (QWQ-32B) and a preprocessing buffer layer, reflecting the multi-model feature fusion and dynamic weight allocation process; 3) Deep Verification and Optimization Layer: Demonstrates the parallel deployment solution and knowledge-enhanced decision-making module for the high-precision DeepSeek model, including an intelligent cache management system (cross-model KV cache sharing, memory pressure prediction) and adversarial cache optimization components (CacheGAN, neural cache topology); 4) Service input and output: Unify the API interface and transmission mechanism of the service gateway.
[0075] Figure 2 : Self-review and weakly supervised and unsupervised labeling using YOLOv8 as an example in this system.
[0076] Figure 3 :An example of the working mechanism of the collaborative verification layer.
[0077] 1) Visual model A (such as Qwen-VL-32B): Outputs a probability map of the lesion area; 2) Visual model D (such as Grad-CAM / YOLOv11): Generates heat maps and locates key lesion areas; 3) Consistency Fusion Module: Calculates pixel-level error using the ICP algorithm and triggers a review mechanism (when the error exceeds the threshold); 4) Reasoning channel: Align the two split visual task inverse fusion reports with the task text cross-modally, constrain the model's attention distribution, and output deep thinking reasoning.
[0078] Figure 4 :An example of the system’s reasoning enhancement process.
[0079] Figure 5 : An example of a three-stage enhancement process.
[0080] Figure 6 : Gateway flow chart.
[0081] 1) Input processing: parsing client requests (images, voice, text); 2) Multimodal processing: Call Qwen2.5-VL to generate image descriptions, which are then combined with user questions and passed to QWQ-32B; 3) Dynamic routing: triggering different processing paths (conventional process / deep analysis channel) based on confidence thresholds or physiological data anomalies; 4) Output integration: Generate diagnostic reports, decision trees or visual interface elements.
[0082] Figure 7 : Adversarial cache optimization design.
[0083] 1) CacheGAN component: The generator predicts cache demand, and the discriminator evaluates the value of cache items; 2) Neural cache topology: semantic graph cache structure and GNN association management; 3) Memory replay mechanism: Periodically replay edge cache content and dynamically update weight coefficients.
[0084] Figure 8 : Enter the test content.
[0085] Figure 9 : After receiving the different feedback reports from the two multimodal models, the reasoning model outputs after deep thinking.
[0086] The above figures collectively illustrate the system architecture, core technology implementation path and collaborative working mechanism between modules of the present invention, providing a clear visual reference for the implementation of the present invention.
Claims
1. A multimodal dynamic image analysis system, characterized in that include: Task decomposition and reconstruction (including but not limited to implementation by embedded small model SLM or direct user input), multimodal or target detection vision module (e.g., a multimodal large model LLM and a target detection model, or, for example, a forward analysis multimodal LLM module and a different target elimination LLM verification module), reasoning enhancement module; And the system performs: a) The task decomposition and reconstruction module decomposes or heterogeneously reconstructs the input multimodal instruction execution task, including but not limited to decomposing it into forward detection tasks, reverse target elimination verification tasks, and fine-grained reasoning enhancement tasks through embedded model reasoning, or decomposing it into target area detection, target attribute analysis, and reasoning enhancement tasks for formation causes; b) The multimodal or object detection vision module performs heterogeneous processing and fusion, including but not limited to processing the forward detection task through a high-precision multimodal model and outputting structured analysis results, while the object detection module generates a verification heat map according to the heterogeneous elimination verification task, performs reverse object elimination, and performs spatial consistency comparison with the forward results, and then performs fusion through an embedded low-precision multimodal model; c) The reasoning enhancement module then performs enhancement work on the analysis result reports of each of the above parts based on task decomposition or instruction intent, including but not limited to logical verification, in-depth interpretation, thinking and reasoning, or targeted fine-grained feature enhancement.
2. A system according to any preceding claim, characterised in that The multimodal and target detection vision module includes at least one multimodal large model module and at least one target detection training module (such as YOLOv8). The working steps include: Step S1: The pre-processing module performs task decomposition and instruction intention attention area positioning; Step S2: The target detection module (such as YOLOv8) generates preliminary target detection or positioning results; Step S3: The multimodal LLM verification module ensures the reliability of the results through a heterogeneous verification mechanism, including but not limited to feature extraction and probability prediction for the detection task; Step S4: The reasoning module summarizes the S2 and S3 reports, performs semantic enhancement on the features, or performs logical verification and in-depth interpretation on the detection results.
3. A system according to any preceding claim, characterised in that include: Multimodal large model module, target detection training module (such as YOLOv8), knowledge base related to the target field, and pseudo-label generation module; The system generates pseudo-labels for images in an unsupervised manner and uses a knowledge base to optimize model training: a) the multimodal large model module receives image input and generates a pseudo label of the target area in combination with the knowledge base; b) the pseudo-label generation module performs spatial consistency verification and selects reliable labels with a confidence level higher than a certain threshold; c) The target detection training module fuses strong labels and weak labels to optimize the target detection training module (such as YOLOv8) model.
4. A system according to any preceding claim, characterised in that The system also includes a cross-model multimodal cache management system, which is characterized by including: Global cache coordinator, cross-model adaptation layer, cache prediction module and hierarchical storage unit; Monitor the KV cache status of each model in real time and implement cross-model KV cache sharing and dynamic compression, and establish a cross-model KV cache collaborative workflow between multimodal models, including but not limited to the following steps: a) Extracting key-value (KV) caches of each attention layer from multiple heterogeneous multimodal models; b) Projecting the KV caches of features of different dimensions into a unified vector space through a cross-model adaptation layer, or mapping the cache spaces of different models into a unified memory pool through hash indexing, such as allowing Qwen-VL-72B and Gemma-27B to share the intermediate results of the visual feature encoding layer; c) Analyze access patterns based on the cache prediction module; d) Implement a tiered storage strategy based on the prediction results (e.g., high-frequency hotspot cache resides in GPU memory, medium-frequency cache is stored in CPU-GPU shared memory, and low-frequency cache is compressed and stored in a solid-state drive array, etc.).
5. A system according to any preceding claim, characterised in that include: Multimodal analysis and target detection modules (e.g., deploying the QwenVL model and YOLOv8, respectively), argument interpretation modules (e.g., deploying the QWQ-32B model), and decision optimization modules (e.g., deploying the DeepSeek 671B model). The system performs: Step S1: The multimodal analysis module generates primary analysis results for the input multimodal data, such as submitting the analysis results and the original input instructions to the reasoning layer argument interpretation module, which then conducts further analysis in combination with the knowledge base. Step S2: The argument interpretation module performs logical verification and reasoning enhancement, such as constructing an argument matrix that includes semantic enhancement and logical verification chains. It generates preliminary diagnosis or event analysis conclusions through multimodal feature fusion and automatically triggers subsequent verification processes based on confidence thresholds. Step S3: The decision optimization module reviews the multimodal analysis results (such as confidence levels lower than 95% or triggering sensitive vocabulary) and generates enhanced conclusions.
6. A system according to any preceding claim, characterised in that It includes the collaborative positioning mechanism of the target detection module (such as YOLOv11) and the multimodal large model. The implementation steps are as follows: Fast region detection: Deploy the target detection module to perform high-precision target detection (such as lung nodules, tumors, and lesion areas) on images, and output the bounding box coordinates (Bbox), category probability, and key point location information of the lesion area; Feature region extraction (optional): Based on the detection results of the target detection module, the system automatically crops or enhances the image blocks of the target area (such as applying CLAHE histogram equalization to low-contrast areas) to generate high-resolution image sub-images; Attention guidance optimization: The detection results of the target detection module are used as the visual attention prior of the multimodal model (such as Qwen-VL-72B). The coordinate information of the lesion area is injected into the model through position encoding (such as Sine Positional Encoding) in the Transformer encoding layer, forcing it to focus on key locations. Joint reasoning process: After receiving the enhanced regional features and detection confidence, the multimodal model dynamically adjusts its contextual understanding weight (for example, in chest X-ray diagnosis, it assigns a double weight to the detected lung shadow area) to generate a more accurate diagnostic conclusion.
7. A system according to any preceding claim, characterised in that The intelligent cache management system includes a dynamic transfer learning module to retrain low-confidence cache data.
8. A system according to any preceding claim, characterised in that To address the attention drift problem of multi-image input, the system adopts a serialization processing mechanism, forcing the multimodal model to parse images one by one, and attaching structured metadata to each image to anchor the context.
9. A system according to any preceding claim, characterised in that The deep parsing layer includes a model parallel deployment solution, which uses parameter sharding technology to decompose large models into multiple logical units (for example, DeepSeek is divided into 8 logical units) to achieve memory affinity optimization under a dual-core NUMA architecture.
10. A system according to any preceding claim, characterised in that The system comprises: (a) Multi-model collaborative architecture: A chain processing pipeline consisting of a preprocessing module, a high-precision multimodal analysis module, a low-precision reverse verification module, a fast inference module, and an enhanced decision module; (b) Task reconstruction mechanism: The input image task is decomposed into a forward detection task and a reverse complementary verification task, which are assigned to different precision models for parallel processing; (c) Dynamic scheduling system for heterogeneous computing resources: Dual-processor and multi-graphics card cluster based on NUMA architecture achieves load balancing; (d) Cross-model knowledge fusion module: Enhanced RAG architecture. The multimodal LLM module retrieves dynamic knowledge index units when analyzing images to achieve retrieval-generation collaboration and realizes multi-model collaborative optimization through methods including but not limited to feature space alignment.
Citation Information
Cited By
Language reasoning server, method for language reasoning, system for visual language big model reasoning, medium and product
CN120725153A
Language inference servers, methods for language inference, systems, media, and products for large-scale visual language model inference.
CN120725153B
Massive network live broadcast batch data acquisition method and system
CN120769077A