A dynamic self-organizing on-board parallel computing method

By adopting star network structure and dynamic arbitration technology in the star-mounted multiprocessor system, the problems of abnormal node detection and recovery are solved, and the computing resources are automatically adjusted by dynamically evaluating the computing power requirements of the algorithm, thereby achieving an efficient, flexible and reliable computing system.

CN115712588BActive Publication Date: 2025-05-23SHANDONG INST OF AEROSPACE ELECTRONICS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211532474.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2025-05-23
Estimated Expiration
2042-12-01

AI Technical Summary

Technical Problem

The prior art cannot automatically detect and eliminate abnormal computing nodes, cannot automatically recover without affecting the integrity of the computing, and cannot independently evaluate the computing power requirements of the algorithm, and automatically increase or decrease computing nodes.

Method used

High-speed bus and switching chips are used to form a star network structure, and data flow is organized through master-slave and division-set methods, master-slave nodes are dynamically arbitrated, abnormal nodes are discovered using transmission timeouts and error judgments, and self-organized through dynamic arbitration requests. At the same time, by counting the calculation time at the main node, dynamically evaluate the computing power requirements of the algorithm, and dynamically increase or decrease the computing nodes.

Benefits of technology

It realizes dynamic self-organization of the computing system, automatically detects and recovers abnormal nodes, and independently adjusts computing resources, improving the efficiency, flexibility and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115712588B_ABST
    Figure CN115712588B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of onboard multi-processor resource scheduling, and in particular to a dynamically self-organizing onboard parallel computing method. A high-speed bus is used in combination with a switching chip to form a star network structure with homogeneous multi-processors; a master-slave and distributed-diversity method is used to organize data streams, with the master node distributing computing instructions or computing data, and the slave nodes collecting the results to the master node after completing their respective calculations; a broadcast ID and a shared overall mapping table are used to autonomously arbitrate master and slave nodes, and re-arbitration is supported at any time, that is, dynamic self-organization is achieved; abnormal nodes are discovered by using a data transmission timeout and error judgment method, and overall re-arbitration is triggered after the abnormality is discovered, so that any node abnormality can be autonomously restored, and no node redundant backup is required; a method of statistically calculating time at the master node and comparing it with the expected value is used to dynamically evaluate the computing power requirements of the algorithm, and dynamically increase or decrease computing nodes to reduce power consumption when the computing power requirements are not high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of satellite-borne multi-processor resource scheduling, and in particular to a satellite-borne parallel computing method capable of dynamic self-organization. Background Art

[0002] With the continuous development of domestic satellite remote sensing technology and satellite constellation technology in recent years, the demand for onboard high-performance computing has continued to increase. The traditional method of transmitting remote sensing data to the ground for processing can no longer meet the needs of real-time and rapid response. At the same time, with the continuous expansion of satellite constellations, satellites need autonomous mission planning to carry out multi-satellite collaborative work, and remote sensing data must be processed in real time to generate intelligence as the basis for autonomous planning.

[0003] With the rapid development of domestic high-performance processors, satellite-borne multi-processor high-performance computing platforms have become possible. However, there are few applications of satellite-borne multi-processor resource scheduling systems in China. If the traditional serial method is used to divide the multi-processors into different tasks, the short board effect will inevitably occur, and the computing resources of the multi-processors will not be able to be fully utilized. The traditional parallel computing method is a fixed organizational form, and the number of computing nodes cannot be dynamically adjusted according to the computing power requirements. In addition, once a computing node is abnormal, the entire computing system will be paralyzed. Summary of the invention

[0004] The present invention provides a dynamically self-organizing onboard parallel computing method, which aims to solve the problem that the prior art cannot automatically detect and eliminate abnormal computing nodes, and cannot automatically recover without affecting the integrity of the computing;

[0005] In addition, the purpose of the present invention is to solve the problem that the prior art cannot autonomously evaluate the computing power requirements of the algorithm and automatically increase or decrease computing nodes.

[0006] To achieve the above object, the technical solution of the present invention is:

[0007] The present invention provides a dynamically self-organizing on-board parallel computing method, comprising:

[0008] At least two high-performance processors with high-speed interfaces, a high-speed bus matching the high-speed interfaces of the high-performance processors, at least one switching chip matching the high-speed bus, and a power module; each of the high-performance processors is connected to the switching chip via a high-speed bus to form a star network structure; the power module is connected to each of the high-performance processors via a low-speed bus, and the power module can individually control the power supply of each of the high-performance processors;

[0009] Assign a unique and fixed ID to each computing node, and dynamically arbitrate the master node and slave node by broadcasting the ID and sharing the overall mapping table for arbitration; the master node distributes computing instructions and computing data, and the slave nodes gather the results to the master node after completing their own calculations; by judging whether there is a transmission timeout or error during each transmission of computing instructions or computing data in a distributed manner, it is determined whether there is an abnormal node, and the abnormal node broadcasts a dynamic arbitration request to the entire system for self-organization; the method of counting the computing time at the master node and comparing it with the expected value is adopted to dynamically evaluate the computing power requirements of the algorithm and dynamically increase or decrease computing nodes;

[0010] The master node distributes computing instructions and computing data, and the slave nodes complete their respective computing and aggregate the results to the master node, including:

[0011] S101. The master node receives a calculation instruction from an external source and determines the calculation algorithm for this calculation;

[0012] S102. The master node receives the data to be processed from the outside and caches it into the memory;

[0013] S103. The master node obtains the number N of slave nodes in the computing network and the ID of each slave node from the overall mapping table;

[0014] S104. The master node distributes computing instructions to all slave nodes;

[0015] S105. The slave node enters the corresponding algorithm and waits to receive data to be processed;

[0016] S106. The master node divides the data to be processed into N parts equally, and distributes one part of the data to be processed to each slave node in turn;

[0017] S107. After receiving a piece of data to be processed from the slave node, the calculation begins;

[0018] S108. After the slave nodes complete the calculation, they send their respective calculation results to the master node;

[0019] S109. After the master node receives the calculation results of all slave nodes in sequence, it integrates them into the final calculation result;

[0020] S110. The master node sends the final calculation result to the outside.

[0021] In one embodiment, the method for arbitration through broadcast ID and shared overall mapping table comprises the following specific steps:

[0022] S201. The computing node broadcasts an arbitration request;

[0023] S202. After receiving the arbitration request, other computing nodes delay for t1 time to filter out duplicate arbitration requests;

[0024] S203. After time t1 is up, each computing node broadcasts its own ID and delays for t2 to wait for all computing nodes to complete broadcasting their own IDs;

[0025] S204. The computing node sorts the received IDs in ascending order to form a mapping table;

[0026] S205. The computing node determines whether the first ID in the mapping table is equal to its own ID. If so, it is the master node, otherwise it is a slave node;

[0027] S206. The master node distributes the overall mapping table to all slave nodes, and the slave nodes receive the overall mapping table from the master node.

[0028] In one embodiment, the step of determining whether an abnormal node occurs by determining whether a transmission timeout or error occurs during each transmission of a computing instruction or computing data in a diversity mode, and the abnormal node broadcasting a dynamic arbitration request to the entire system for self-organization includes: sending a determination program and receiving a determination program;

[0029] The sending determination procedure comprises:

[0030] S301. The transmitting end sends a short frame to each receiving end, wherein the short frame includes the length, type and label information of the calculation instruction or calculation data;

[0031] S302. The sender determines whether the sending of the short frame has timed out or the length is wrong;

[0032] S303. If there is a failure and it exceeds N times, it is considered abnormal and the sender broadcasts an arbitration request;

[0033] S304. If normal, receive a response from the receiving end;

[0034] S305. The sender determines whether the response has timed out, whether the length is correct, and whether the response content is correct;

[0035] S306. If the response fails and exceeds N times, it is considered abnormal and an arbitration request is broadcast;

[0036] S307. If the response is normal, the calculation instruction or calculation data is sent;

[0037] S308. Determine whether the calculation instruction or calculation data sent has timed out or has a length error;

[0038] S309. If the failure occurs more than N times, it is considered abnormal and an arbitration request is broadcast;

[0039] S310. If normal, receive a response;

[0040] S311. Determine whether the response has timed out, whether the length is correct, and whether the response content is correct;

[0041] S312. If the response fails and exceeds N times, it is considered abnormal and an arbitration request is broadcast;

[0042] S313. If the response is normal, the transmission is completed;

[0043] The receiving determination procedure comprises:

[0044] S320. The receiving end receives the short frame;

[0045] S321. Determine whether the received short frame has timed out or has a length error;

[0046] S322. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast;

[0047] S323. If normal, compare the length, type, and label of the short frame;

[0048] S324. The comparison result forms a response frame and sends a response frame;

[0049] S325. Determine whether the response timeout has occurred and whether the length is correct;

[0050] S326. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast;

[0051] S327. If the response is normal, the calculation instruction or calculation data is received;

[0052] S328. Determine whether the received calculation instruction or calculation data has timed out or has a length error;

[0053] S329. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast;

[0054] S330. If normal, send a response;

[0055] S331. Determine whether the response timeout has occurred and whether the length is correct;

[0056] S332. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast;

[0057] S333. If the sending response is normal, the reception is completed.

[0058] In one embodiment, the step of dynamically evaluating the computing power requirements of the algorithm and dynamically increasing or decreasing computing nodes by using a method of counting computing time at the master node and comparing it with an expected value includes:

[0059] S401. The master node counts the time at the beginning and end of each calculation, and obtains the duration of a calculation, and compares it with the preset maximum expected value or minimum expected value;

[0060] S402. If the calculation time of a certain algorithm is greater than the maximum expected value for n consecutive times, the computing power requirement is calculated; the number of computing nodes that need to be awakened = (computation time - expected value) ÷ expected value × the number of current nodes; the master node powers on the required idle nodes, and after powering on, the idle nodes will issue a dynamic arbitration request to join the mapping table;

[0061] S403. If the calculation time of a certain algorithm is less than the minimum expected value for n consecutive times, the computing power surplus is calculated: the number of computing nodes that need to be shut down = (expected value - calculation time) ÷ expected value × current number of nodes; the master node cuts off power to the redundant nodes in order of ID from large to small, and then initiates a dynamic arbitration request to re-form the mapping table.

[0062] The beneficial effects achieved by the present invention are:

[0063] The present invention adopts a high-speed bus combined with a switching chip to form a star network structure with homogeneous multi-processors; adopts a master-slave and distribution-diversity mode to organize data flow, that is, one processor is used as a master node, and other processors are used as slave nodes. The master node distributes calculation instructions or calculation data, and the slave nodes gather the results to the master node after completing their own calculations; adopts a broadcast ID and a shared overall mapping table mode to autonomously arbitrate master and slave nodes, and supports re-arbitration at any time, that is, to achieve dynamic self-organization; adopts a data transmission timeout and error judgment mode to discover abnormal nodes, and triggers overall re-arbitration after the abnormality is discovered, so that any node abnormality can be autonomously restored, and no node redundant backup is required; adopts a method of counting the calculation time at the master node and comparing it with the expected value, dynamically evaluates the computing power requirement of the algorithm, and dynamically increases or decreases the computing nodes to achieve reduced power consumption when the computing power requirement is not high. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.

[0065] Figure 1 It is a schematic diagram of the overall architecture of the onboard parallel computing system.

[0066] Figure 2 It is a processing flow chart of the master-slave node of the present invention.

[0067] Figure 3 It is a flow chart of the dynamic arbitration of the present invention.

[0068] Figure 4 The present invention is a flow chart of abnormal node discovery and recovery based on transmission timeout and error judgment.

[0069] Figure 5 It is a flowchart of autonomous computing power evaluation and adjustment based on the comparison of computing time and expected value of the present invention.

[0070] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0071] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0072] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0073] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, if the meaning of "and / or" appearing in the full text is to include three parallel schemes, taking "A and / or B" as an example, it includes scheme A, or scheme B, or a scheme that satisfies both A and B. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of ordinary technicians in this field to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0074] The present invention provides a dynamically self-organizing satellite-borne parallel computing method, which organizes data streams in a master-slave and distribution-diversity manner, can automatically detect abnormal computing nodes, and can automatically recover without affecting the integrity of the calculation; in addition, it can also autonomously evaluate the computing power requirements of the algorithm, automatically increase or decrease computing nodes, and has power consumption control capabilities, so that the computing system has high efficiency, flexibility and reliability.

[0075] The onboard parallel computing method is based on a star network structure composed of homogeneous multi-processors, such as Figure 1 As shown, the star network structure composed of the homogeneous multi-processors includes at least two high-performance processors with high-speed interfaces, a high-speed bus matching the high-speed interfaces of the high-performance processors, at least one switching chip matching the high-speed bus, and a power module; each of the high-performance processors is connected to the switching chip through a high-speed bus to form a star network structure; the power module is connected to each of the high-performance processors through a low-speed bus, and the power module can control the power supply of each of the high-performance processors individually. Each high-performance processor is a computing node. After forming a star network, any two computing nodes can communicate at high speed through the switching chip to transmit computing data information.

[0076] Among them, the high-speed interface can be selected from SRIO, PCIE or 10 Gigabit Ethernet, etc.; the high-performance processor can be selected from domestic FT-M6678, domestic FT-D2000 or PowerPC-P2020, etc.; the low-speed bus can be selected from RS422 or CAN bus, etc.

[0077] A processor (ie, the master node described below) can control the power on and off of other processors through the power module as needed.

[0078] If the number of interfaces of a switching chip is insufficient, multiple switching chips can be cascaded; each switching chip is connected to several high-performance processors through a high-speed bus, and the switching chips are also connected by a high-speed bus, forming a star network.

[0079] The power supply module is connected to each of the high-performance processors through a low-speed bus to provide power to each of the high-performance processors. The power supply module can control the power on and off of other processors based on the control of a certain high-performance processor (i.e., the main node described below); the power supply module can directly adopt existing technology, so this article does not describe its structure and principle in detail.

[0080] Assign a unique and fixed ID to each computing node, and dynamically arbitrate the master node and slave node by broadcasting the ID and sharing the overall mapping table for arbitration, that is, one processor is the master node and the other processors are slave nodes. The master node distributes computing instructions and computing data, and the slave nodes gather the results to the master node after completing their own calculations. At the same time, by judging whether there is a transmission timeout or error during each transmission of computing instructions or computing data in a distributed-diversity manner, it is determined whether there is an abnormal node and the abnormal node ID is determined. The abnormal node broadcasts a dynamic arbitration request to the entire system for self-organization. At the same time, the method of counting the computing time at the master node and comparing it with the expected value is adopted to dynamically evaluate the computing power requirements of the algorithm, and dynamically increase or decrease computing nodes to achieve reduced power consumption when the computing power requirements are not high.

[0081] like Figure 2 As shown, the master node distributes computing instructions and computing data, and the slave nodes complete their respective calculations and aggregate the results to the master node, including:

[0082] S101. The master node receives a calculation instruction from an external source and determines the calculation algorithm for this calculation;

[0083] S102. The master node receives the data to be processed from the outside and caches it into the memory;

[0084] S103. The master node obtains the number N of slave nodes in the computing network and the ID of each slave node from the overall mapping table;

[0085] S104. The master node distributes computing instructions to all slave nodes;

[0086] S105. The slave node enters the corresponding algorithm and waits to receive data to be processed;

[0087] S106. The master node divides the data to be processed into N parts equally, and distributes one part of the data to be processed to each slave node in turn;

[0088] S107. After receiving a piece of data to be processed from the slave node, the calculation begins;

[0089] S108. After the slave nodes complete the calculation, they send their respective calculation results to the master node;

[0090] S109. After the master node receives the calculation results of all slave nodes in sequence, it integrates them into the final calculation result;

[0091] S110. The master node sends the final calculation result to the outside.

[0092] This method allows only the master node of the entire computing system to communicate externally, hiding the details of the computing system and simplifying the design of the entire satellite or constellation system. This method evenly divides the calculation so that the algorithm does not need to specify the number of calculation blocks, but dynamically obtains the number of slave nodes; thus, the algorithm applied to this system does not need to be bound by the number of processors as in the traditional serial or parallel method, and the system flexibility is improved.

[0093] like Figure 3 As shown, the method for arbitration through broadcast ID and shared overall mapping table comprises the following specific steps:

[0094] S201. The computing node broadcasts an arbitration request;

[0095] S202. After receiving the arbitration request, other computing nodes delay for t1 time to filter out duplicate arbitration requests;

[0096] S203. After time t1 is up, each computing node broadcasts its own ID and delays for t2 to wait for all computing nodes to complete broadcasting their own IDs;

[0097] S204. The computing node sorts the received IDs in ascending order to form a mapping table;

[0098] S205. The computing node determines whether the first ID in the mapping table is equal to its own ID. If so, it is the master node, otherwise it is a slave node;

[0099] S206. The master node distributes the overall mapping table to all slave nodes, and the slave nodes receive the overall mapping table from the master node.

[0100] Each node broadcasts its own ID to all nodes, and each node forms an overall ID mapping table, which determines the master and slave nodes in order of ID size. However, a node may be abnormal, resulting in a failure to broadcast its ID or a failure to receive an ID. Therefore, the master node needs to distribute the overall mapping table to all slave nodes to ensure that the IDs held by all computing nodes are consistent and unique. Among them, for computing nodes that fail to broadcast, this node will be excluded when forming the mapping table. After confirming the master node, the master node will broadcast its own ID, and the master node and slave nodes will store this ID.

[0101] The way to trigger dynamic arbitration is: a computing node broadcasts an arbitration request. When the computing node is powered on or an abnormality is found, the arbitration request will be broadcast.

[0102] The introduction of the dynamic arbitration method makes the computing system a flexible computing system, which can dynamically increase or decrease computing nodes according to the needs of computing power, and dynamically restore or eliminate abnormal nodes, thereby ensuring system reliability while improving flexibility.

[0103] like Figure 4 As shown, the step of determining whether an abnormal node appears and determining the abnormal node ID by determining whether a transmission timeout or error occurs during each transmission of a computing instruction or computing data in a distributed manner, and the abnormal node broadcasting a dynamic arbitration request to the entire system for self-organization includes a sending determination program and a receiving determination program;

[0104] The sending decision procedure of the sending end (i.e., the computing node that sends the computing instruction or computing data) includes the following steps:

[0105] S301. The transmitting end sends a short frame to each receiving end, wherein the short frame includes the length, type and label information of the calculation instruction or calculation data;

[0106] S302. The sender determines whether the sending of the short frame has timed out or the length is wrong;

[0107] S303. If there is a failure and it exceeds N times, it is considered abnormal and the sender broadcasts an arbitration request;

[0108] S304. If normal, receive a response from the receiving end;

[0109] S305. The sender determines whether the response has timed out, whether the length is correct, and whether the response content is correct;

[0110] S306. If the response fails and exceeds N times, it is considered abnormal and an arbitration request is broadcast;

[0111] S307. If the response is normal, the calculation instruction or calculation data is sent;

[0112] S308. Determine whether the calculation instruction or calculation data sent has timed out or has a length error;

[0113] S309. If the failure occurs more than N times, it is considered abnormal and an arbitration request is broadcast;

[0114] S310. If normal, receive a response;

[0115] S311. Determine whether the response has timed out, whether the length is correct, and whether the response content is correct;

[0116] S312. If the response fails and exceeds N times, it is considered abnormal and an arbitration request is broadcast;

[0117] S313. If the response is normal, the transmission is completed.

[0118] The receiving decision procedure of the receiving end (i.e., the computing node that receives the computing instruction or computing data) includes the following steps:

[0119] S320. The receiving end receives the short frame;

[0120] S321. Determine whether the received short frame has timed out or has a length error;

[0121] S322. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast;

[0122] S323. If normal, compare the length, type, and label of the short frame;

[0123] S324. The comparison result forms a response frame and sends a response frame;

[0124] S325. Determine whether the response timeout has occurred and whether the length is correct;

[0125] S326. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast;

[0126] S327. If the response is normal, the calculation instruction or calculation data is received;

[0127] S328. Determine whether the received calculation instruction or calculation data has timed out or has a length error;

[0128] S329. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast;

[0129] S330. If normal, send a response;

[0130] S331. Determine whether the response timeout has occurred and whether the length is correct;

[0131] S332. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast;

[0132] S333. If the sending response is normal, the reception is completed.

[0133] By judging whether a transmission timeout or error occurs during each transmission of a computing instruction or computing data in a distributed manner, it is determined whether an abnormal node occurs and the abnormal node ID is determined.

[0134] The basic operations of the distribution-diversity transmission are sending and receiving. At the same time, the master and slave nodes in the system perform opposite operations, that is, when the master node is sending, the slave node is receiving; and when the slave node is sending, the master node is receiving. In addition, since the calculation starting point is when the master node distributes the calculation command, the calculation starting points of all nodes are basically the same. Each node calculates the equally divided data blocks of the data to be processed, and the calculation time is basically the same. In this way, for a certain calculation, the operation timing in the system can be basically estimated, and the maximum waiting time for each transmission or reception can be estimated.

[0135] Among them, the waiting time for the master node to send to each slave node is: the number of bytes to be sent × transmission rate; the waiting time for each slave node to receive the master node is: the number of bytes to be received × transmission rate × the number of slave nodes n (the last node must wait for the previous nodes to complete); the waiting time for each slave node to send to the master node is: the number of bytes to be sent × transmission rate × the number of slave nodes n (there may be slave nodes that send at the same time and need to wait); the waiting time for the master node to receive each slave node is: the number of bytes to be received × transmission rate. In order to leave a certain margin, a time t is added to each waiting time. When a certain sending or receiving operation times out, it can be considered that a node abnormality has occurred (it may be an abnormality on this side or the other side).

[0136] While judging the transmission timeout, the correctness of the transmitted data is judged. Since the master and slave nodes send and receive operations in pairs, that is, each send and receive knows the data type and length, before the actual valid data, the sender first sends a short data frame composed of type, length, and specific message tags, and the receiver confirms the short data frame. After confirming that it is correct, the receiver replies with a response. If there is no correct response, it can be judged as abnormal (it may be abnormal on this side or abnormal on the other side). At the same time, the actual length of each send or receive is judged. If it is an unexpected value, it is judged as abnormal.

[0137] When a node is abnormal due to timeout or data correctness judgment, it broadcasts a dynamic arbitration request to the entire system for self-organization. Since abnormal nodes cannot enter dynamic arbitration, the re-formed mapping table will not include abnormal nodes. After self-organization, the master node will power off the abnormal node and choose to power on other idle nodes according to the computing power requirements. After the idle node is powered on, it will initiate a dynamic arbitration request to join the mapping table. At the same time, the node will record the number of abnormal nodes as a basis for whether to enable them again in the future.

[0138] After recovering from the exception, the master node determines whether its own ID has changed compared with the stored master node ID. If it has not changed, it will redo the calculation interrupted by the exception; if it has changed, the cached data to be processed will be lost, and the master node needs to request the outside world to resend the data to be processed.

[0139] Compared with the traditional heartbeat judgment method, this method does not require redundant transmission and communication. When an abnormal node is found, the master and slave nodes can be re-determined by triggering arbitration to form an overall mapping table, and the abnormality can be recovered and the calculation can be continued in time, ensuring the integrity of the calculation. Compared with the traditional hot backup method, this method does not require backup and saves hardware resources.

[0140] like Figure 5 As shown, the method of using the method of counting the computing time on the master node and comparing it with the expected value to dynamically evaluate the computing power requirements of the algorithm and dynamically increase or decrease the computing nodes includes:

[0141] S401. The master node counts the time at the beginning and end of each calculation, and obtains the duration of one calculation, and compares it with the preset maximum expected value or minimum expected value; the maximum expected value or minimum expected value is estimated based on the algorithm type and data length, and is a known value that can be expected;

[0142] S402. If the calculation time of a certain algorithm is greater than the maximum expected value for n consecutive times, the computing power requirement is calculated, that is, how many nodes need to be woken up again; the number of computing nodes that need to be woken up again = (calculation time - expected value) ÷ expected value × the current number of nodes; the master node powers on the required idle nodes, and after powering on, the idle nodes will issue a dynamic arbitration request to join the mapping table.

[0143] S403. If the calculation time of a certain algorithm is less than the minimum expected value for n consecutive times, the computing power margin is calculated, that is, how many nodes can be shut down: the number of computing nodes to be shut down = (expected value - calculation time) ÷ expected value × current number of nodes. The master node powers off the redundant nodes in descending order of ID, and then initiates a dynamic arbitration request to re-form the mapping table.

[0144] This method can dynamically evaluate the computing power requirements of the algorithm and dynamically increase or decrease computing nodes to reduce power consumption when the computing power requirements are not high.

[0145] Specifically, the present invention is described in detail using the following examples:

[0146] 1. Computing Network Composition

[0147] The high-performance processor uses the domestic high-performance DSP - Feiteng FT-M6678N, with a total of 16 DSPs, each with a main frequency of 1GHz and a DDR3 capacity of 2GB. The high-speed bus uses SRIO, with a single-channel rate of 5Gbps. The SRIO switching chip uses two RXS2448s, and each switching chip connects to 8 DSPs. The switching chip and one DSP use one SRIOx1, and one SRIOx4 is used between the switching chips. At the same time, one switching chip is connected to the external host computer through two SRIOx4s.

[0148] The low-speed bus uses RS422, and each DSP is connected to the power supply module through RS422. The power supply module supplies power to 16 DSPs. Each DSP can control the power supply module to power on or off other DSPs through RS422.

[0149] 2. Data Flow Organization

[0150] The specific steps of the data flow group are as follows:

[0151] S101. Assign fixed IDs to 16 DSPs: 0 to 15;

[0152] S102. The master node receives a calculation instruction from an external source and determines the calculation algorithm for this calculation;

[0153] S103. The master node receives the data to be processed from the outside and caches it into memory;

[0154] S104. The master node obtains the number N (N is less than or equal to 16) of slave nodes in the computing network and the ID of each slave node from the overall mapping table;

[0155] S105. The master node distributes computing instructions to all slave nodes;

[0156] S106. The slave node enters the corresponding algorithm and waits to receive data to be processed;

[0157] S107. The master node divides the data to be processed into N parts equally, and distributes one part of the data to be processed to each slave node in turn;

[0158] S108. After receiving a piece of data to be processed from the slave node, the calculation begins;

[0159] S109. After the slave nodes complete the calculation, they send their respective calculation results to the master node;

[0160] S110. After the master node receives the calculation results of all slave nodes in sequence, it integrates them into the final calculation result;

[0161] S111. The master node sends the final calculation result to the outside.

[0162] 3. Dynamic Node Arbitration

[0163] The specific steps of dynamic node arbitration are as follows:

[0164] S201. The computing node broadcasts an arbitration request;

[0165] S202. After receiving the arbitration request, other computing nodes delay 200ms to filter out duplicate arbitration requests;

[0166] S203. After the time is up, each computing node broadcasts its own ID and delays for 200ms to wait for all nodes to complete broadcasting their own IDs;

[0167] S204. The computing node sorts the received IDs in ascending order to form a mapping table;

[0168] S205. The computing node determines whether the first ID in the mapping table is equal to its own ID. If so, it is the master node, otherwise it is a slave node;

[0169] S206. The master node distributes the mapping table to all slave nodes, and the slave nodes receive the mapping table from the master node.

[0170] 4. Abnormal Node Discovery and Recovery

[0171] The specific steps for abnormal node discovery and recovery are as follows:

[0172] Sending side:

[0173] S301. Group the short frames by length, type and tag and send them out;

[0174] S302. Determine whether the short frame has timed out or has a length error;

[0175] S303. If the failure occurs more than 3 times, it is considered abnormal and an arbitration request is broadcast;

[0176] S304. If normal, receive a response;

[0177] S305. Determine whether the response has timed out, whether the length is correct, and whether the response content is correct;

[0178] S306. If the response fails and exceeds 3 times, it is considered abnormal and an arbitration request is broadcast;

[0179] S307. If the response is normal, the data frame is sent;

[0180] S308. Determine whether the data frame has timed out or has a length error;

[0181] S309. If the failure occurs more than 3 times, it is considered abnormal and an arbitration request is broadcast;

[0182] S310. If normal, receive a response;

[0183] S311. Determine whether the response has timed out, whether the length is correct, and whether the response content is correct;

[0184] S312. If the response fails and exceeds 3 times, it is considered abnormal and an arbitration request is broadcast;

[0185] S313. If the response is normal, the transmission is completed.

[0186] Receiver:

[0187] S320. Receive short frames;

[0188] S321. Determine whether the received short frame has timed out or has a length error;

[0189] S322. If the failure occurs more than 3 times, it is considered abnormal and an arbitration request is broadcast;

[0190] S323. If normal, compare the length, type, and label of the short frame;

[0191] S324. The comparison result forms a response frame and sends a response frame;

[0192] S325. Determine whether the response timeout has occurred and whether the length is correct;

[0193] S326. If the failure occurs more than 3 times, it is considered abnormal and an arbitration request is broadcast;

[0194] S327. If the response is normal, the data frame is received;

[0195] S328. Determine whether the received frame has timed out or has a length error;

[0196] S329. If the failure occurs more than 3 times, it is considered abnormal and an arbitration request is broadcast;

[0197] S330. If normal, send a response;

[0198] S331. Determine whether the response timeout has occurred and whether the length is correct;

[0199] S332. If the failure occurs more than 3 times, it is considered abnormal and an arbitration request is broadcast;

[0200] S333. If the sending response is normal, the reception is completed.

[0201] 5. Autonomous computing power evaluation and adjustment

[0202] The specific steps for autonomous computing power evaluation and adjustment are as follows:

[0203] S401. The master node calculates the starting point t1;

[0204] S402. Start a calculation;

[0205] S403. The master node calculates the end point time t2;

[0206] S404. Obtain a calculation duration t3 = t2 - t1;

[0207] S405. Obtaining the preset maximum expected value and minimum expected value of the calculation time according to the algorithm type;

[0208] S406. Determine whether the duration is greater than the maximum expected value for 5 consecutive times;

[0209] S407. If yes, calculate the number of nodes added = (calculation time - maximum expected value) ÷ maximum expected value × current number of nodes, power on the newly added nodes, and perform dynamic arbitration;

[0210] S408. If not, continue to determine whether the duration is less than the minimum expected value for 5 consecutive times;

[0211] S409. If yes, then calculate the number of redundant nodes = (minimum expected value - calculation time) ÷ minimum expected value × current number of nodes, power off the redundant nodes, and perform dynamic arbitration;

[0212] S410. If no, this calculation ends.

[0213] The above descriptions are only optional embodiments of the present invention, and are not intended to limit the patent scope of the present invention. All equivalent structural changes made using the contents of the present invention's specification and drawings, or directly / indirectly applied in other related technical fields, are included in the patent protection scope of the present invention.

Claims

1. A dynamically self-organizing on-board parallel computing method. It is characterized in that include: At least two high-performance processors with high-speed interfaces, a high-speed bus matching the high-speed interfaces of the high-performance processors, at least one switching chip matching the high-speed bus, and a power module; each of the high-performance processors is connected to the switching chip via a high-speed bus to form a star network structure; The power supply module is connected to each of the high-performance processors through a low-speed bus, and the power supply module can individually control the power supply of each of the high-performance processors; Assign a unique and fixed ID to each computing node, and dynamically arbitrate the master node and slave node by broadcasting the ID and sharing the overall mapping table for arbitration; the master node distributes computing instructions and computing data, and the slave nodes gather the results to the master node after completing their own calculations; by judging whether a transmission timeout or error occurs during each transmission of computing instructions or computing data in a distributed-diversity manner, it is determined whether an abnormal node exists, and the abnormal node broadcasts a dynamic arbitration request to the entire system for self-organization; By using the method of counting the computing time on the master node and comparing it with the expected value, the algorithm's computing power requirements are dynamically evaluated, and computing nodes are dynamically increased or decreased; The master node distributes computing instructions and computing data, and the slave nodes complete their own calculations and aggregate the results to the master node, including: S101. The master node receives a calculation instruction from an external source and determines the calculation algorithm for this calculation; S102. The master node receives the data to be processed from the outside and caches it into memory; S103. The master node obtains the number N of slave nodes in the computing network and the ID of each slave node from the overall mapping table; S104. The master node distributes computing instructions to all slave nodes; S105. The slave node enters the corresponding algorithm and waits to receive data to be processed; S106. The master node divides the data to be processed into N parts equally, and distributes one part of the data to be processed to each slave node in turn; S107. After receiving a piece of data to be processed from the slave node, the calculation begins; S108. After the slave nodes complete the calculation, they send their respective calculation results to the master node; S109. After the master node receives the calculation results of all slave nodes in sequence, it integrates them into the final calculation result; S110. The master node sends the final calculation result to the outside.

2. A dynamically self-organizing on-board parallel computing method according to claim 1, Features: The method for arbitration by broadcasting ID and sharing the overall mapping table includes the following specific steps: S201. The computing node broadcasts an arbitration request; S202. After receiving the arbitration request, other computing nodes delay for t1 time to filter out duplicate arbitration requests; S203. After time t1 is up, each computing node broadcasts its own ID and delays for t2 to wait for all computing nodes to complete broadcasting their own IDs; S204. The computing node sorts the received IDs in ascending order to form a mapping table; S205. The computing node determines whether the first ID in the mapping table is equal to its own ID. If so, it is the master node, otherwise it is a slave node; S206. The master node distributes the overall mapping table to all slave nodes, and the slave nodes receive the overall mapping table from the master node.

3. The method for dynamically self-organizing on-board parallel computing according to claim 1, Features: By judging whether a transmission timeout or error occurs during each transmission of computing instructions or computing data in a distributed-diversity manner, whether an abnormal node occurs, and the abnormal node broadcasts a dynamic arbitration request to the entire system for self-organization, including: sending a determination program and receiving a determination program; The sending determination procedure comprises: S301. The transmitting end sends a short frame to each receiving end, wherein the short frame includes the length, type and label information of the calculation instruction or calculation data; S302. The sender determines whether the sending of the short frame has timed out or the length is wrong; S303. If there is a failure and it exceeds N times, it is considered abnormal and the sender broadcasts an arbitration request; S304. If normal, receive a response from the receiving end; S305. The sender determines whether the response has timed out, whether the length is correct, and whether the response content is correct; S306. If the response fails and exceeds N times, it is considered abnormal and an arbitration request is broadcast; S307. If the response is normal, the calculation instruction or calculation data is sent; S308. Determine whether the calculation instruction or calculation data sent has timed out or has a length error; S309. If the failure occurs more than N times, it is considered abnormal and an arbitration request is broadcast; S310. If normal, receive a response; S311. Determine whether the response has timed out, whether the length is correct, and whether the response content is correct; S312. If the response fails and exceeds N times, it is considered abnormal and an arbitration request is broadcast; S313. If the response is normal, the transmission is completed; The receiving determination procedure comprises: S320. The receiving end receives the short frame; S321. Determine whether the received short frame has timed out or has a length error; S322. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast; S323. If normal, compare the length, type, and label of the short frame; S324. The comparison result forms a response frame and sends a response frame; S325. Determine whether the response timeout has occurred and whether the length is correct; S326. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast; S327. If the response is normal, the calculation instruction or calculation data is received; S328. Determine whether the received calculation instruction or calculation data has timed out or has a length error; S329. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast; S330. If normal, send a response; S331. Determine whether the response timeout has occurred and whether the length is correct; S332. If the failure occurs more than M times, it is considered abnormal and an arbitration request is broadcast; S333. If the sending response is normal, the reception is completed.

4. The method for dynamically self-organizing on-board parallel computing according to claim 1, Features: The method of statistically calculating the computing time on the master node and comparing it with the expected value is used to dynamically evaluate the computing power requirements of the algorithm and dynamically increase or decrease the steps of computing nodes, including: S401. The master node counts the time at the beginning and end of each calculation, and obtains the duration of a calculation, and compares it with the preset maximum expected value or minimum expected value; S402. If the calculation time of a certain algorithm is greater than the maximum expected value for n consecutive times, the computing power requirement is calculated; the number of computing nodes that need to be awakened = (computation time - expected value) ÷ expected value × the number of current nodes; the master node powers on the required idle nodes, and after powering on, the idle nodes will issue a dynamic arbitration request to join the mapping table; S403. If the calculation time of a certain algorithm is less than the minimum expected value for n consecutive times, the computing power surplus is calculated: the number of computing nodes that need to be shut down = (expected value - calculation time) ÷ expected value × current number of nodes; the master node cuts off power to the redundant nodes in order of ID from large to small, and then initiates a dynamic arbitration request to re-form the mapping table.

Citation Information

Patent Citations

  • Method and device for processing system-on-chip shared bus requests

    CN103257942A

  • Commercial low-earth-orbit satellite communication system

    CN113131990A