System architecture supporting distributed synchronous computing

By introducing components such as multi-core DSP processors, synchronization control modules and dynamic load balancing modules into distributed computing systems, the problems of device coordination and clock differences are solved, efficient and accurate data transmission and system synchronization are achieved, and the overall performance and security of the system are improved.

CN120256370AActive Publication Date: 2025-07-04BEIJING ASTRONAUTICS JUHENG SYST INTEGRATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510749263.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing distributed computing systems have shortcomings in data synchronization, real-time and resource utilization efficiency. The equipment coordination of motion axis-related computing and control tasks is poor, the data transmission lacks effective rules, clock differences lead to calculation errors, and the lack of an effective clock calibration mechanism.

Method used

The multi-core DSP processor in the motion control device group is interconnected through the SRIO bus, the single-board computer configures the system timing and shared data rules, the exchange card sets up and downlink data interfaces, the synchronization control module has 5 synchronization signals and priority judgment, the dynamic load balancing module adopts machine learning algorithm, data transit across core direct connection channels, the motion control device internal clock calibration, the single-board computer performs data prefetching, and the exchange card performs error detection and correction.

Benefits of technology

It improves the system operation efficiency, ensures accurate and orderly transmission of data between devices, avoids calculation errors, realizes coordinated and synchronous operation of various parts of the system, protects data security, avoids network congestion and improves data transmission accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256370A_ABST
    Figure CN120256370A_ABST
Patent Text Reader

Abstract

The invention discloses a system architecture supporting distributed synchronous calculation, which comprises a motion control equipment group, a single-board computer, an exchange card, a synchronous control module, a dynamic load balancing module, a cross-core direct connection channel and the like, wherein the motion control equipment group comprises a multi-core DSP (Digital Signal Processor), is interconnected through an SRIO (Serial Radio Input / Output) bus, and has downlink directional distribution and error detection and correction functions. The synchronous control module has five paths of synchronous signals which have different functions, and a priority judgment and queuing mechanism is arranged in the synchronous control module. The single-board computer is provided with a predictive data prefetching mechanism and an encryption chip. The cross-core direct connection channel is provided with a specific process. The machine learning algorithm of the dynamic load balancing module has a self-adaptive updating mechanism. The local clock of the motion control device can be calibrated. The method has the advantages of improving operation efficiency, guaranteeing accurate and orderly transmission of data, coordinating synchronous operation of the system, protecting data safety, avoiding network congestion, improving data transmission accuracy and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of application of computing system architectures, and particularly to a system architecture supporting distributed synchronous computing. Background Art

[0002] With the rapid development of information technology, distributed computing has been widely applied in many fields, such as big data processing, cloud computing, artificial intelligence, etc. However, existing distributed computing systems still have some deficiencies in aspects such as data synchronization, real-time performance, and resource utilization efficiency. Traditional distributed systems often rely on centralized control methods, resulting in limitations in the scalability and flexibility of the system. In addition, delays and packet losses may occur during data transmission, affecting the overall computing performance.

[0003] In common distributed synchronous computing system architectures, the device collaboration for motion axis-related calculations and control tasks is poor, resulting in low overall operating efficiency of the system. At the same time, there is a long waiting time during the data processing process, and there is a lack of effective rules for data transmission between devices, which is prone to chaos and errors, affecting the normal operation of the system. In common system architectures, there is also a lack of effective control and management mechanisms for synchronization signals, and it is impossible to reasonably arrange the processing order of synchronization signals with different functions, thus affecting the coordinated synchronous operation of various parts of the system. In addition, it is difficult to keep the clocks of devices consistent, and clock differences are prone to cause calculation errors. There is a lack of an effective clock calibration mechanism and it cannot meet the working requirements of the application of the computing system architecture. Therefore, a system architecture supporting distributed synchronous computing is proposed. Summary of the Invention

[0004] The present invention provides the following technical solutions: A system architecture supporting distributed synchronous computing, comprising: A motion control device group, a single-board computer, a switching card, a synchronization control module, a dynamic load balancing module, and a cross-core direct connection channel. Each motion control device in the motion control device group is internally equipped with a multi-core DSP processor, and the multi-core DSP processor is used for distributed motion axis control. The motion control devices in the motion control device group are interconnected through an SRIO bus; The single-board computer configures the system timing and shared data rules through a PCIE interface. The switching card includes an upstream data interface, a downstream data interface, and a shared data interface. The internal of the synchronization control module is provided with 5 synchronization signals, including Sync1 - Sync5. The internal of the dynamic load balancing module is provided with a machine learning algorithm; Each motion control device in the motion control device group is internally equipped with a local clock, and the internal of the single-board computer is provided with a predictive data prefetching mechanism.

[0005] Preferably, when receiving data from downstream devices, the upstream data interface packs the data according to the "checkerboard" rule and broadcasts it to all motion control devices via SRIO. The downstream data interface distributes the motion control device data to downstream devices according to the routing ID.

[0006] Preferably, Sync1 in the synchronization control module is used to trigger the DSP interrupt of the motion control device, execute calculations and data distribution. Sync2 in the synchronization control module is used to control the upstream and downstream data distribution timing of the switch card. Sync3 in the synchronization control module is used to trigger downstream device synchronization through a doorbell. Sync4 in the synchronization control module is used to control the broadcast of shared data. Sync5 in the synchronization control module is used to convert to a PCIE interrupt packet to trigger the task scheduling of the single-board computer.

[0007] Preferably, an encryption chip is installed inside the single-board computer. When the single-board computer configures the system timing and shared data rules through the PCIE interface, data encryption is performed through the encryption chip and the AES algorithm inside it.

[0008] Preferably, the establishment of the cross-core direct connection channel includes the source motion control device marking the SRIO-ID of the target motion control device in the data header, then the switch card resolves the ID for data transfer, and finally the channel is automatically released after the transmission is completed.

[0009] Preferably, the upstream data interface of the switch card is set with a flow control function and adopts the token bucket algorithm to allocate data traffic according to the receiving capacity of the upper-layer device. The token bucket algorithm sets the rate of token generation and the capacity of the bucket. Only when there are enough tokens can the data be transmitted upward through the upstream data interface.

[0010] Preferably, the downstream data interface of the switch card is internally set with an error detection and correction function. The error detection and correction function adopts cyclic redundancy check and Hamming code technology, which can detect whether there are errors in the data at the data receiving end and correct the error data.

[0011] Preferably, the local clock inside the motion control device has a clock calibration function. Through the clock calibration function, and using the network time protocol and GPS clock signal for calibration, the local clocks are kept consistent.

[0012] Preferably, the machine learning algorithm inside the dynamic load balancing module is set with an adaptive update mechanism. By collecting the load data of the system, the load data includes the task execution time and resource utilization rate of each device. The machine learning algorithm can adjust its own parameters and model structure according to these load data.

[0013] Preferably, the synchronization control module needs to have a priority judgment and queuing mechanism inside. The synchronization control module sets priorities for the 5 synchronization signals through its internal priority judgment and queuing mechanism, and can sort and process different synchronization signals according to preset priority rules.

[0014] In summary, compared with the prior art, the present invention provides a system architecture supporting distributed synchronous computing, which has the following beneficial effects: 1. The present invention uses a multi-core DSP processor in a motion control device group for distributed motion axis control, and the devices are interconnected through an SRIO bus, which helps to efficiently perform motion axis-related calculation and control tasks, realize collaborative work between devices, and improve the overall operating efficiency of the system. In addition, through the predictive data pre-fetching mechanism of the single-board computer, it is possible to obtain data that may be used in advance, reduce data waiting time, and increase the speed of data processing, thereby speeding up the operation rhythm of the entire system. At the same time, the uplink data interface of the switching card packages and broadcasts data according to the Tianzi grid rule, and the downlink data interface distributes data in a direction according to the routing ID. This regularized data processing method ensures that data is transmitted accurately and orderly between different devices, reducing confusion and errors in data transmission; 2. The present invention uses 5 synchronization signals inside the synchronization control module to respectively undertake different functions such as triggering interrupts, controlling timing, triggering downstream device synchronization, controlling shared data broadcasting, and triggering single-board computer task scheduling. It also has an internal priority judgment and queuing mechanism, which can reasonably arrange the processing order of the synchronization signals to ensure the coordinated and synchronous operation of various parts of the system. The local clock inside the motion control device has a calibration function, and the network time protocol and GPS clock signal are combined to keep each local clock consistent, providing a unified time reference for the distributed synchronous calculation of the system, ensuring the time synchronization of different devices, and avoiding calculation errors caused by clock differences. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic diagram of the system architecture structure of the present invention.

[0016] Figure 2 It is a schematic diagram of the structure of the switching card of the present invention.

[0017] Figure 3 It is a structural schematic diagram of the synchronous control module of the present invention. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] Please refer to Figure 1 , the present invention provides a system architecture supporting distributed synchronous computing, including: A group of motion control devices, a single-board computer, a switch card, a synchronization control module, a dynamic load balancing module, and a cross-core direct connection channel. Multi-core DSP processors are installed inside the motion control devices in the group of motion control devices, and the multi-core DSP processors are used for distributed motion axis control. The motion control devices in the group of motion control devices are interconnected through SRIO buses; A single-board computer, which configures the system timing and shared data rules through a PCIE interface. Please refer to Figure 2 , the switch card includes an upstream data interface, a downstream data interface, and a shared data interface. When the upstream data interface receives data from downstream devices, it packs the data according to the grid rule and broadcasts it to all motion control devices through SRIO. The specific process of packing according to the grid rule is as follows; Data reception preparation: The upstream data interface of the switch card is in a state of waiting to receive data from downstream devices and is ready to perform a packing operation according to the grid rule; Data reception and grouping. The upstream data interface sequentially receives data units from downstream devices. These data units may be original data segments sent in a certain order. According to the grid rule, the received data is grouped. Assuming that the data is regarded as a continuous information flow, it is divided into four parts in the shape of a grid with a fixed number of bytes or data block size as the unit (this is only a conceptual division, and it may be different in actual according to specific data structures and protocols). For example, every 4 consecutive data units are divided into a group. The first and fourth data units are used as the upper left and lower right parts of the grid, and the second and third data units are used as the upper right and lower left parts of the grid; Packing operation: Pack each group of data divided according to the grid rule. This may involve adding some identification information to identify the grouping structure and order of the data during the subsequent unpacking process. For example, a specific identifier is added at the beginning of each group of data to indicate that this is a data group packed according to the grid rule, and the order information of the data within the group is recorded. The packed data groups are combined together to form a complete data packet for subsequent broadcasting to all motion control devices through SRIO; The downstream data interface distributes the motion control device data to downstream devices according to the routing ID. And there is an error detection and correction function inside the downstream data interface of the switching card. The error detection and correction function adopts cyclic redundancy check and Hamming code technology, which can detect whether there are errors in the data at the data receiving end and correct the error data. The specific process of the above data error detection and correction is as follows: Data reception preparation: The downstream data interface of the switching card distributes the motion control device data to downstream devices according to the routing ID. When the data arrives at the data receiving end of the downstream data interface, the system is ready to start the error detection and correction function, which adopts cyclic redundancy check and Hamming code technology. Error detection step: First, using cyclic redundancy check technology, the system calculates the received data according to a pre-set generating polynomial to obtain a CRC check code. Subsequently, the calculated CRC check code is compared with the CRC check code attached to the received data. If the two check codes are the same, it is initially judged that the data may not have been in error during transmission; if they are different, it indicates that the data may be in error. If the CRC detects that the data may be in error, then the Hamming code technology is started for further detection. According to the coding rules of the Hamming code, the data is analyzed to determine the location of the error (if there is an error). Error correction step: If the Hamming code detects an error and can determine the error location, the error data bit is corrected according to the error correction rules of the Hamming code. The data after correction will be regarded as correct data and continue to be transmitted or processed in the system. If the CRC detection indicates that the data may be correct or no error is found after the Hamming code detection, then the data will directly follow the normal process for subsequent transmission, processing, etc. in the system. Please refer to Figure 3 , there are 5-way synchronization signals inside the synchronization control module, including Sync1 - Sync5. Sync1 in the synchronization control module is used to trigger the DSP interruption of the motion control device to execute calculation and data distribution. Sync2 in the synchronization control module is used to control the upstream and downstream data distribution timing of the switching card. Sync3 in the synchronization control module is used to trigger downstream device synchronization through a doorbell. Sync4 in the synchronization control module is used to control the shared data broadcast. Sync5 in the synchronization control module is used to convert to a PCIE interruption packet to trigger the task scheduling of the single-board computer. There is a machine learning algorithm inside the dynamic load balancing module. The machine learning algorithm inside the dynamic load balancing module is provided with an adaptive update mechanism. By collecting the load data of the system, where the load data includes the task execution time and resource utilization rate of each device, the machine learning algorithm can adjust its own parameters and model structure according to these load data. The specific process of the above steps is as follows; Load data collection: The system determines the range of devices for which load data needs to be monitored, covering various devices in the distributed synchronous computing system architecture, such as motion control device groups, single-board computers, switch cards, etc. For each device, it starts collecting load data. Among them, the task execution time is obtained by recording the time taken from when the device starts executing a task to when the task is completed. In terms of resource utilization, for example, for a multi-core DSP processor, monitor the usage ratios of resources such as its CPU usage rate and memory occupancy rate, and collect and summarize these data such as task execution time and resource utilization rate as load data; Data processing and analysis: Preprocess the collected load data. For example, remove outliers (such as unreasonable data caused by temporary device failures or sudden interferences), and perform data formatting to make it meet the input requirements of machine learning algorithms. According to the characteristics of machine learning algorithms, specific features may be extracted from the load data. For example, extract features related to task types from the task execution time data to better analyze the impact of different types of tasks on the system load. Input the processed load data into the machine learning algorithm, and the algorithm evaluates the current load status of the system based on the existing model structure and parameters, determining whether the system load is balanced, which devices have excessive or insufficient loads, etc.; Parameter and model structure adjustment: If the evaluation result shows that the system load is unbalanced, the machine learning algorithm decides how to adjust its own parameters according to the predefined strategy. For example, if the task execution time of a certain device is too long and the resource utilization rate is too high, the algorithm may decide to adjust the weight parameter of the task assigned to this device. In some cases, if the current model structure cannot well adapt to the load change situation, the algorithm will optimize the model structure. For example, increase or decrease the hidden layer, adjust the neuron connection method, etc. According to the new parameters and model structure, reassign tasks or adjust the resource allocation strategy to achieve the balance of the system load and improve the overall performance of the system. As the system runs, continuously repeat the above processes of load data collection, analysis, and adjustment, enabling the machine learning algorithm to continuously adapt to the dynamic changes of the system load and continuously optimize the load balancing effect of the system; Local clocks are installed inside the motion control devices in the motion control device group, and a predictive data prefetch mechanism is set inside the single-board computer. The specific process of the predictive data prefetch mechanism is as follows; Data requirement analysis: The single-board computer monitors the tasks it is currently executing and about to execute. This includes obtaining task-related information, such as task types and task priorities, through interactions with other components (such as the synchronization control module, dynamic load balancing module, etc.). According to the task information, analyze the data requirements associated with these tasks. For example, if a task involves operating a certain device in the motion control device group, determine the data that may need to be obtained from this device or other related devices, such as device status data, historical operation record data, etc.; Prefetch trigger judgment: Evaluate the current resource status of the single-board computer, including memory capacity, processing power, etc. Ensure that there are sufficient resources to perform the prefetch operation without affecting the currently ongoing tasks. Based on the results of data demand analysis and resource evaluation, determine whether the prefetch conditions are met. The prefetch conditions may include the importance of the data (such as data crucial for upcoming high-priority tasks), the acquisition cost of the data (such as low cost for obtaining from local cache and high cost for obtaining from remote devices), etc. If the prefetch conditions are met, for example, there is sufficient memory space and the data is very important for the upcoming task, then trigger the prefetch operation; Data prefetch execution: Determine the data source from which the data needs to be prefetched. This may be data in other motion control devices, switch cards, or local storage. Send a prefetch request to the data source. If the data source is other devices, send the request through the corresponding interface (such as PCIE interface or SRIO bus). The request contains information identifying the data to be prefetched, such as the address and type of the data. Receive the data returned from the data source and cache it in the local cache of the single-board computer. During the caching process, some preprocessing may be performed on the data, such as data format conversion, etc., so that subsequent tasks can use it quickly; Prefetch data update and management: Regularly check the validity of the prefetched data. Since the system state may change, the prefetched data may become outdated. For example, if the state of a certain device changes, the previously prefetched state data of that device may no longer be accurate. If the prefetched data is invalid, trigger the prefetch operation again according to the new requirements and update the data in the local cache. Or, if there are new data requirements and there is sufficient space in the local cache, new data can also be directly prefetched and part of the old data can be replaced; The local clock inside the motion control device has a clock calibration function. Through the clock calibration function, and using the Network Time Protocol and GPS clock signal for calibration, the local clocks are kept consistent. The specific steps of the above process are as follows; Clock calibration preparation: The local clock inside each motion control device in the motion control device group is initialized at device startup and starts timing. The predictive data prefetch mechanism of the single-board computer runs normally, independent of but simultaneous with the local clock calibration process; Clock calibration signal acquisition: The motion control device sends a request to the Network Time Protocol server through the network connection to obtain the NTP signal. This NTP signal contains accurate time information and is used as one of the references for calibrating the local clock. For motion control devices with GPS reception function, receive the GPS clock signal through the GPS receiver. The GPS clock signal also carries high-precision time data and can be used for local clock calibration; Clock Calibration Operation: The motion control device compares the current time of its local clock with the time in the received NTP signal and GPS clock signal (if available), calculates the time deviation between the local clock and the standard time (NTP or GPS time), and based on the calculated time deviation, the clock calibration function inside the motion control device starts to adjust the local clock. If the local clock is faster than the standard time, the time of the local clock is slowed down to match the standard time by reducing the clock frequency or other appropriate methods. If the local clock is slower than the standard time, the clock frequency is increased or similar adjustment means are adopted to make the time of the local clock catch up with the standard time. The above comparison and adjustment steps are continuously repeated to ensure that the local clocks inside each motion control device in the entire motion control device group are consistent. In this way, in the distributed synchronous computing system architecture, each motion control device can operate based on a consistent local clock, avoiding calculation and control errors caused by clock differences; An encryption chip is installed inside the single-board computer. When the single-board computer configures the system timing and shared data rules through the PCIE interface, data encryption is performed through the encryption chip and the AES algorithm inside it. The specific process of the above steps is as follows; Data Preparation: The single-board computer obtains the data related to configuring the system timing and shared data rules through the PCIE interface. This data may include various system parameters, device identifiers, and other information; Encryption Preparation: The encryption chip inside the single-board computer performs an initialization operation before data encryption, loads necessary configuration information, such as key length, encryption mode, and other related settings, initializes the AES algorithm inside the encryption chip, and determines the specific operation mode of encryption according to the preset standards (such as AES - 128, AES - 192, or AES - 256, etc.), for example, Electronic Codebook (ECB) mode, Cipher Block Chaining (CBC) mode, etc.; Encryption Operation: According to the requirements of the AES algorithm, if the data length exceeds the block size specified by the algorithm (such as 128 bits), the data is divided into several fixed-size blocks, and the encryption chip encrypts each data block according to the AES algorithm. During the encryption process, the data block is encrypted and transformed using the pre-set key (this key is securely stored inside the encryption chip). For each data block, after multiple rounds of encryption operations (such as AES - 128 may require 10 rounds of encryption operations), the encrypted data block is obtained; Processing of Encrypted Data: The single-board computer uses the encrypted data for the configuration operations of the system timing and shared data rules. When these encrypted data are transmitted and processed inside the system, the security and privacy of the data can be guaranteed, preventing the data from being stolen or tampered with during the configuration process; The establishment of the cross-core direct connection channel includes the source motion control device marking the SRIO-ID of the target motion control device in the data packet header, then the switch card parses the ID for data transfer, and finally automatically releases the channel after the transmission is completed; The upstream data interface of the switch card is set with a flow control function and adopts the token bucket algorithm. According to the receiving capacity of the upper-layer device, the data flow is allocated. The token bucket algorithm sets the rate of token generation and the capacity of the bucket. Only when there are enough tokens can the data be transmitted upward through the upstream data interface. The specific process of the above steps is as follows; Token bucket initialization: Determine two key parameters of the token bucket according to the receiving capacity of the upper-layer device, including the rate of token generation (r) and the capacity of the bucket (C). For example, if the upper-layer device can receive 100 data units per second, the token generation rate r can be set to 100 tokens per second, and the capacity of the bucket C is set to 200 tokens (assumed value) according to factors such as the system cache capacity; Token generation and accumulation: The token bucket continuously generates tokens at the set token generation rate r. For example, tokens are added to the bucket at the rate of r per second, and the generated tokens accumulate in the bucket but do not exceed the capacity C of the bucket. If the bucket is full, the newly generated tokens will be discarded; Data transmission determination: When data arrives at the upstream data interface of the switch card and is ready to be transmitted to the upper-layer device, flow control determination is required to check whether there are enough tokens in the token bucket to allow data transmission. If the data size is n data units, it is necessary to check whether there are at least n tokens in the bucket. For example, if a data packet requires 5 tokens and there are only 3 tokens in the bucket, then the data packet cannot be transmitted temporarily; Data transmission and token consumption: If there are enough tokens in the bucket (the number of tokens is greater than or equal to the number of tokens required for the data), the data is allowed to be transmitted upward through the upstream data interface to the upper-layer device. When transmitting data, the corresponding number of tokens is consumed according to the data size. For example, after transmitting a data packet that requires 5 tokens, the number of tokens in the bucket is reduced by 5; Continuous monitoring and adjustment: Continuously monitor the data flow situation through the upstream data interface, including the arrival rate and transmission volume of the data. According to the changes in the system operation situation and the receiving capacity of the upper-layer device, it may be necessary to adjust the token generation rate r and the capacity of the bucket C. For example, if the receiving capacity of the upper-layer device increases, the value of r can be appropriately increased to allow more data transmission; The synchronization control module needs to have a priority judgment and queuing mechanism inside. The synchronization control module sets the priorities of the 5-way synchronization signals through its internal priority judgment and queuing mechanism, and can sort and process different synchronization signals according to the preset priority rules. The specific process of the above steps is as follows; Priority rule determination: Based on the system architecture and the functions of each synchronization signal, determine the importance of each synchronization signal during system operation. For example, Sync1 is used to trigger the DSP interrupt of the motion control device, perform calculations, and distribute data, which is crucial for the normal operation of the motion control device; Sync2 controls the uplink and downlink data distribution timing of the switch card, affecting the order of data transmission, etc. According to the functional importance, preset priority rules are formulated. Assuming that the priority order is set according to the degree of influence on the key functions of the system as: Sync1>Sync2>Sync3>Sync4>Sync5; Priority setting: The synchronization control module identifies 5 synchronization signals: Sync1 - Sync5, and marks the corresponding priority for each synchronization signal according to the preset priority rules. For example, mark the highest priority for Sync1 and the lowest priority for Sync5; Queuing mechanism: When a synchronization signal arrives at the synchronization control module, it enters the corresponding queue according to its priority mark. High-priority signals enter the head of the queue, and low-priority signals enter the tail of the queue. For example, when Sync1 arrives, it will be at the head of the queue, and when Sync5 arrives, it will be at the tail of the queue; Sorting and processing: The synchronization control module processes the synchronization signals sequentially from the head according to the order in the queue, and first processes the high-priority synchronization signals. For example, first process Sync1 to perform its functions of triggering the DSP interrupt of the motion control device, performing calculations, and distributing data. During the process of processing high-priority signals, if a low-priority signal arrives, the low-priority signal needs to wait until the high-priority signal is processed and then be processed in order. For example, when processing Sync1, if Sync5 arrives, Sync5 needs to wait until Sync1 is processed, and then process Sync5 according to the queue order (convert Sync5 into a PCIE interrupt packet to trigger the task scheduling of the single-board computer). During system operation, if the system state changes, for example, the task requirements of some devices change or a fault occurs, it may be necessary to dynamically adjust the priority rules, re-evaluate the importance of each synchronization signal according to the new system state, adjust the priority order, and accordingly re-mark the priorities and re-queue, and then continue to process the synchronization signals in the new order.

[0020] In this solution, a multi-core DSP processor in the motion control device group is used for distributed motion axis control, and the devices are interconnected through an SRIO bus. This helps to efficiently perform motion axis-related calculation and control tasks, realize collaborative work among devices, improve the overall operation efficiency of the system, and through the predictive data prefetch mechanism of the single-board computer, it is able to prefetch data that may be used in advance, reduce data waiting time, and enhance the speed of data processing, thereby accelerating the operation rhythm of the entire system. At the same time, the upstream data interface of the switch card packs data according to the grid rule and broadcasts it, and the downstream data interface distributes data directionally according to the routing ID. This regular data processing method ensures the accurate and orderly transmission of data between different devices, reducing chaos and errors in data transmission.

[0021] In this solution, 5 synchronization signals inside the synchronization control module are used to respectively undertake different functions such as triggering interrupts, controlling timing, triggering downstream device synchronization, controlling shared data broadcasting, and triggering single-board computer task scheduling. And there is a priority judgment and queuing mechanism inside, which can reasonably arrange the processing order of synchronization signals to ensure the coordinated and synchronous operation of all parts of the system. Moreover, the local clock inside the motion control device has a calibration function, and combined with the network time protocol and GPS clock signal, each local clock is kept consistent, providing a unified time reference for the distributed synchronous calculation of the system, ensuring the synchronization of different devices in time, and avoiding calculation errors caused by clock differences.

[0022] When this solution configures the system timing and shared data rules in the single-board computer, data encryption is performed through an internal encryption chip and the AES algorithm to effectively protect the security and privacy of data, preventing data from being stolen or tampered with during the configuration process. And the upstream data interface of the switch card has a flow control function, and the token bucket algorithm is used to allocate traffic according to the receiving capacity of the upper-layer device to avoid network congestion caused by excessive data traffic and ensure the stability of data transmission; the error detection and correction function of the downstream data interface uses cyclic redundancy check and Hamming code technology to detect and correct data errors in a timely manner, improving the accuracy of data transmission.

[0023] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0024] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A system architecture for supporting distributed synchronous computing, characterized in that, Including: A group of motion control devices, a single-board computer, a switch card, a synchronization control module, a dynamic load balancing module, and a cross-core direct connection channel. Multi-core DSP processors are installed inside the motion control devices in the group of motion control devices, and the multi-core DSP processors are used for distributed motion axis control. The motion control devices in the group of motion control devices are interconnected through SRIO buses. A single-board computer configures system timing and shared data rules through a PCIE interface. The switch card includes an upstream data interface, a downstream data interface, and a shared data interface. Five synchronization signals, including Sync1 - Sync5, are set inside the synchronization control module. A machine learning algorithm is set inside the dynamic load balancing module. Local clocks are installed inside the motion control devices in the group of motion control devices, and a predictive data prefetch mechanism is set inside the single-board computer.

2. The system architecture for supporting distributed synchronous computing according to claim 1, characterized in that: When the upstream data interface receives data from downstream devices, it packs the data according to the grid rule and broadcasts it to all motion control devices through SRIO. The downstream data interface distributes the motion control device data to downstream devices directionally according to the routing ID.

3. A system architecture for supporting distributed synchronous computing according to claim 1, characterized in that: Sync1 in the synchronization control module is used to trigger the DSP interrupt of the motion control device, execute calculations and data distribution. Sync2 in the synchronization control module is used to control the upstream and downstream data distribution timing of the switch card. Sync3 in the synchronization control module is used to trigger the synchronization of downstream devices through a doorbell. Sync4 in the synchronization control module is used to control shared data broadcasting. Sync5 in the synchronization control module is used to convert to a PCIE interrupt packet to trigger the task scheduling of the single-board computer.

4. A system architecture for supporting distributed synchronous computing according to claim 1, characterized in that: An encryption chip is installed inside the single-board computer. When the single-board computer configures system timing and shared data rules through the PCIE interface, data encryption is performed through the encryption chip and the AES algorithm inside it.

5. A system architecture for supporting distributed synchronous computing according to claim 1, characterized in that: The establishment of the cross-core direct connection channel includes the source motion control device marking the SRIO-ID of the target motion control device in the data header, then the switch card resolves the ID for data transfer, and finally the channel is automatically released after the transmission is completed.

6. The system architecture for supporting distributed synchronous computing according to claim 1, characterized in that: The upstream data interface of the switch card is set with a flow control function and adopts the token bucket algorithm to allocate data traffic according to the receiving capacity of the upper-layer device. The token bucket algorithm sets the rate of token generation and the capacity of the bucket. Only when there are enough tokens can the data be transmitted upward through the upstream data interface.

7. The system architecture for supporting distributed synchronous computing according to claim 1, wherein: The downstream data interface of the switch card is internally set with an error detection and correction function. The error detection and correction function adopts cyclic redundancy check and Hamming code technology to detect whether there are errors in the data at the data receiving end and correct the error data.

8. A system architecture for supporting distributed synchronous computing according to claim 1, characterized in that: The local clock inside the motion control device has a clock calibration function. Through the clock calibration function, and using the network time protocol and GPS clock signal for calibration, all local clocks are kept consistent.

9. The system architecture for supporting distributed synchronous computing according to claim 1, characterized in that: The machine learning algorithm inside the dynamic load balancing module is set with an adaptive update mechanism. By collecting the load data of the system, where the load data includes the task execution time and resource utilization rate of each device, the machine learning algorithm can adjust its own parameters and model structure according to the load data.

10. The system architecture for supporting distributed synchronous computing according to claim 1, wherein: The synchronization control module needs to have a priority judgment and queuing mechanism inside. The synchronization control module sets priorities for 5-way synchronization signals through its internal priority judgment and queuing mechanism, and can sort and process different synchronization signals according to the preset priority rules.

Citation Information

Patent Citations

  • Synchronization method based on distributed-type integrated recorder parallel buses

    CN101882989A

  • VPX bus-based workpiece bench synchronous motion control system and method

    CN105511502A

  • Distributed motion control system based on CMC (Control Module on Chip)

    CN106094741A

  • Uniform load distributing method for use in executing parallel processing in parallel computer

    US5535387A