The invention provides a plastic architecture design and
pruning method based on heterogeneous experts and a related device, and the method comprises the steps: constructing a multi-
modal large model architecture with a multi-layer heterogeneous
expert group, considering task difference and environment heterogeneity, enabling each group of experts to have different scales and calculation overhead, integrating the resource
perception capability into a model training process, and obtaining a multi-
modal large model architecture with a multi-layer heterogeneous
expert group. The model actively perceives parameter resource limitation on the architecture level, on the basis, expert combinations are enumerated on a
calibration set in the post-training
pruning stage, redundant experts with low contribution degree are permanently discarded to reduce static parameter quantity, high-contribution experts are dynamically activated and low-contribution experts are skipped during reasoning according to input features and expert weights in the dynamic
pruning stage, and the static parameter quantity is reduced. And reasoning efficiency is improved. According to the scheme, collaborative optimization is carried out from
model architecture design to the reasoning process, the consumption of
model parameters and computing resources is greatly reduced while the multi-
modal task performance is guaranteed, and an innovative solution is provided for efficient deployment of a large multi-modal model in a resource-constrained environment.