The application provides a multi-
modal large
language model instruction federal fine-tuning
system, method and product, and relates to the technical field of
electronic information. The
system comprises: a
server issuing a plurality of task parameters of a current round to a
client; the
client modulates image tokens using the
semantics of local samples to obtain multi-
modal representation, and then processes the multi-
modal representation using local individualized parameters of the current round to obtain multi-modal individualized representation; and the
server uploads the parameters of the current round obtained by fine-tuning the local model based on the representation to obtain individualized parameters of the current round which are left in the local; the
server aggregates the fine-tuned parameters of the current round of the same task according to the fine-tuned parameters of the current round uploaded by each
client to obtain and issue the next round parameters of the task; after fine-tuning is completed, the client obtains a final local model based on the local final individualized parameters and the final parameters corresponding to the task issued by the server, and uses the final local model for local task execution. The purpose is to improve the generalization performance of collaborative fine-tuning.