This invention provides a method for accelerating real-time
inference of multimodal large models, relating to the field of
data processing technology. The method includes: Step 1, generating a minimum envelope
ellipse using the peak point of the first
visual attention weight, the second visual coordinate point, and the third visual association point as boundary points, and calculating the ratio of the area to the perimeter of the
ellipse as a
spatial clustering index. This invention effectively improves the efficiency, accuracy, and
resource utilization of real-time
inference of multimodal large models by selectively enhancing spatial features, evaluating multidimensional constraint computing power and dynamically scheduling resources, lightweight cross-
modal fusion, and multi-step
inference attention calibration, combined with real-time decision-making and full-process data-driven evaluation parameters and dynamic updates of the model
library. This ensures the real-time performance and reliability of inference results, while continuously optimizing
system computing power evaluation and model
adaptation, making multimodal inference more aligned with the
resource constraints of heterogeneous computing and the real-time and accuracy requirements of actual business operations.