The application belongs to the field of
information processing and is a unified memory allocation management method for AI training of a heterogeneous architecture, comprising the following steps: S1: collecting hardware counter indexes and
software layer operator graph
topology information, calculating the instantaneous performance
power consumption ratio of each heterogeneous node, and generating a global resource atlas with a weight
label; S2: identifying a high communication operator
cluster based on the global resource atlas, mapping the required memory to a physically adjacent area, constructing a cross-
chip virtual continuous
address space, and dynamically migrating the load under the
delay constraint according to the performance
power consumption curve; S3: predicting the future
survival period of a
memory block by using a timing model, triggering lazy fragment compression only for the region predicted to be released, and inserting the arrangement micro-operation into the idle period of the computing core for execution; S4: collecting
throughput,
delay,
energy consumption and fragment arrangement
hit rate data, calculating a reward function, and updating the weight parameters of the timing model in real time, and dynamically switching the special sub-model strategy according to the characteristics of the current training stage.