The invention relates to the field of
data processing, in particular to an
adaptive computing power segmentation scheduling method for
large model reasoning tasks, which comprises the following steps of: synchronously acquiring characteristics of common dimensions of a plurality of tasks to be reasoned and computing power cluster equipment; clustering the tasks, setting a plurality of batch size schemes according to the tasks, and dividing task batches; for each scheme, the tasks are matched with equipment according to the batch sequence, the matching degree of batches is calculated, and the weights are determined by the ratio of the feature weights to the same-dimension features of the equipment tasks; after matching is completed, counting the equipment computing power, the memory
utilization rate and the balance degree, and obtaining a scheme evaluation index in combination with the matching degree; and comparing the scheme evaluation indexes under all the
processing modes, and selecting the highest scheme and the matching result thereof as a final scheme. According to the method, the multi-dimensional features and multiple
processing modes are comprehensively considered, the optimal scheme is selected for task and equipment matching, the task requirement can be better met, the
processing time of the task on the equipment is shortened, and therefore service
delay is reduced.