The invention relates to the technical field of large models, in particular to a
large model distributed reasoning acceleration method based on multi-
modal feature fusion and dynamic weight optimization, and the method comprises the following steps: S1, carrying out real-time semantic analysis through a reasoning context analysis module embedded in a load balancer to obtain semantic features; s2, calculating a cache
adaptation degree according to the semantic features; s3, querying a global cache
directory service to obtain a matched node
list, and making a decision by using a multi-fusion decision
algorithm in combination with various parameters of cache
adaptation condition, node load condition and network quality; and S4, deploying a global cache
directory service in the load balancing layer according to the decision, maintaining a local cache
pool at each computing node, and
synchronizing the cache
metadata to the global cache
directory service in real time. The method solves the problems that when a conventional load balancing strategy is adopted for multi-
node deployment, waste of computing resources is easily caused, the overall
energy consumption of the
system is increased, and the request
processing speed is reduced.