Disclosed in the embodiments of the present application is a
collective communication method. The method is applied to a
collective communication system, wherein the
collective communication system comprises a management node and N computing nodes, the N computing nodes being used for concurrently
processing a first communication task, and N being an integer greater than 1. The method comprises: a management node receiving first information, wherein the first information indicates that at least one first link has failed, and each first link is a
communication link between N computing nodes that is used for
processing the first communication task; the management node receiving N pieces of second information from the N computing nodes, wherein each piece of second information is used for indicating whether
original data of the first communication task that is locally stored in a computing node has been changed; and if the
original data in the N computing nodes has not been changed, the management node sending third information to the N computing nodes, so as to instruct the N computing nodes to re-process the first communication task on the basis of the
original data. In this way, the time required for task
recovery after a communication fault occurs can be reduced.