Multi-axis nanoscale real-time synchronous control system and method based on heterogeneous computing collaboration

The multi-axis nanoscale real-time synchronous control system based on heterogeneous computing collaboration solves the synchronization and computing real-time problems of traditional systems in nanoscale positioning scenarios, improves nanoscale positioning accuracy and computing efficiency, reduces hardware costs and power consumption, and meets the accuracy and efficiency requirements of high-end equipment.

CN120779846AActive Publication Date: 2025-10-14WU XI XING WEI KE JI YOU XIAN GONG SI HANG ZHOU FEN GONG SI

Patent Information

Application Number
CN202511242502.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-14
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Traditional multi-axis motion control systems have difficulty meeting the requirements of synchronization, real-time computing, and high reliability in nano-level positioning scenarios, resulting in serious positioning errors and mechanical resonance problems, and are unable to meet the precision requirements of high-end equipment.

Method used

A multi-axis nanoscale real-time synchronous control system based on heterogeneous computing collaboration is adopted. Through the time-space separation architecture module, global synchronization mechanism and collaborative workflow, the dual DSP cores and FPGA of the AM5728 processor are used to form a computing-intensive layer and a time-sensitive layer. Combined with OCXO clock source, LVDS network, PCIe or GPMC interface, emergency stop protection circuit and other technologies, multi-axis data synchronous acquisition, calculation optimization and output are achieved.

Benefits of technology

It achieves nanosecond-level synchronization of multi-axis control signals, with a positioning accuracy of ±7.2 nanometers, improved computing efficiency, reduced hardware costs, reduced power consumption, and improved safety, meeting the precision and efficiency requirements of high-end equipment such as semiconductor manufacturing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120779846A_ABST
    Figure CN120779846A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of high-precision motion control, and discloses a multi-axis nanoscale real-time synchronous control system and method based on heterogeneous computing cooperation. The system adopts a space-time separation architecture of an AM5728 processor and an FPGA (Field Programmable Gate Array); the FPGA is used as a time sensitive layer; nanosecond synchronous acquisition of data of a multi-axis sensor and strict synchronous output of a control signal are realized through a global clock network; a double-DSP core of the AM5728 serves as a calculation dense layer, and parallel calculation of a multi-axis control algorithm is executed. Low-latency interaction of instructions and data is achieved through an optimized PCIe / GPMC interface. Under a clock taming mechanism, the synchronization error of the multi-axis control signal is obviously reduced, and the positioning precision is obviously improved. According to the scheme, the problem of real-time cooperative control of a multi-axis system in the fields of semiconductor manufacturing and the like is solved, and the cost is greatly reduced compared with a traditional architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of precision motion control technology, and in particular to a multi-axis nanoscale real-time synchronous control system based on heterogeneous computing collaboration. Background Art

[0002] In high-end equipment fields such as semiconductor manufacturing and precision optical processing, the performance of multi-axis motion control systems directly affects product precision and production efficiency. For example, the workpiece stage of a photolithography machine must simultaneously control 4-6 precision motion axes to achieve nanometer-level positioning, and the control signals of each axis must be strictly synchronized. Traditional microcontroller-based solutions have fundamental flaws: the characteristic of the microcontroller executing control loops sequentially results in millisecond-level time differences in the output of multi-axis signals. For example, a certain type of photolithography machine has an inter-axis delay of more than 100 microseconds, resulting in a positioning error of 120 nanometers, which cannot meet the requirements of the 3-nanometer process.

[0003] As control algorithms become more complex, computational times for advanced algorithms like model predictive control (MPC) on microcontrollers exceed 200 microseconds, far exceeding the required 50-microsecond control cycle. This forces systems to reduce control frequency or simplify algorithms, sacrificing dynamic response performance. Furthermore, control signal jitter (typically ±50 microseconds) can excite mechanical resonance, accumulating to cause trajectory deviations exceeding 100 nanometers in high-speed motion scenarios.

[0004] Existing improvement proposals attempt to overcome this bottleneck through two approaches: First, using a dedicated motion control chip, which improves computing speed but increases hardware costs by over 40%. Second, employing a field-programmable gate array plus microcontroller architecture, however, the microcontroller's computing power still lacks the necessary support for real-time 16-axis model predictive control. More critically, the software-level emergency stop response time exceeds 1 millisecond, making it impossible to promptly interrupt excessive motion in nanopositioning scenarios, posing a risk of equipment collision.

[0005] The essence of the above problem is that traditional architectures cannot simultaneously meet the triple requirements of nanosecond synchronization, real-time performance of complex algorithms, and high reliability. The current industry urgently needs a new system architecture that can coordinate timing control and computational load, which is the core research direction of this invention. Summary of the Invention

[0006] In order to solve the above technical problems, the present invention provides a multi-axis nanoscale real-time synchronization control system based on heterogeneous computing collaboration, characterized in that the system includes the following modules: Time-space separation architecture module: The dual DSP cores of the AM5728 processor form the computationally intensive layer, and the FPGA forms the time-sensitive layer; Global synchronization mechanism module: OCXO clock source is distributed to FPGA and AM5728 via LVDS network to reduce phase deviation; Collaborative Workflow Module: (a) FPGA synchronously collects multi-axis data and adds time stamps on the rising edge of the clock; (b) Data is transmitted to the DSP via the PCIe or GPMC interface to achieve communication optimization; (c) DSP completes the control algorithm calculation within the specified time; (d) The calculation results are transmitted back to the FPGA and the multi-axis outputs are synchronously updated on the next rising edge of the clock.

[0007] Furthermore, the FPGA includes: A 24-channel ADC synchronous sampling unit, wherein the sampling jitter of the sampling unit is less than a first threshold; 16-axis PWM synchronous output unit, updated by the same clock edge trigger; An emergency stop protection circuit, wherein a response time of the emergency stop protection circuit is less than a second threshold.

[0008] Furthermore, the DSP core runs the TI-RTOS real-time operating system; achieves computational acceleration by adopting vectorized instructions for parallel computing; and integrates a watchdog timer to monitor whether a task cycle has timed out.

[0009] Furthermore, the specific structure of the PCIe or GPMC is: The PCIe protocol frame structure consists of a 12-byte header and a 256-byte payload, with a bandwidth utilization rate of 95%; or The GPMC interface timing parameters are address line setup time 4ns and data line hold time 2ns.

[0010] Furthermore, the FPGA performs nanometer-level interpolation on the encoder's original signal to improve signal resolution, and monitors axis offset in real time through a hardware position comparator; the DSP dynamically identifies the mechanical resonant frequency and injects an anti-phase compensation signal, combined with an adaptive proportional-integral-differential control algorithm to adjust the gain parameters in real time according to changes in load inertia, thereby improving multi-axis positioning accuracy.

[0011] The present invention also provides a multi-axis synchronous control method of a multi-axis nanoscale real-time synchronous control system based on heterogeneous computing collaboration, which is characterized by comprising the following steps: Step S1: Configure the OCXO master clock to trigger full system synchronization; Step S2: Use vectorized instructions to calculate the 16-axis control quantities in parallel:

[0012] Where U is the axis control output, E is the axis error vector, is the proportional gain matrix, is the integral gain matrix, is the differential gain matrix; Step S3: Synchronously update all axis output signals at the rising edge of the next clock cycle.

[0013] The embodiments of the present invention have the following technical effects: The multi-axis synchronous control system of the present invention achieves all-round performance breakthroughs in the field of semiconductor manufacturing. Relying on the hardware-level global clock trigger mechanism and high-precision clock taming network of the field programmable gate array, the synchronization error of the multi-axis control signal output is stably controlled within the nanosecond level, which is many times higher than the microsecond-level error of the traditional microcontroller architecture. It completely eliminates the trajectory distortion problem caused by the inter-axis delay of the lithography machine worktable, and achieves a hundred-fold improvement in synchronization performance. By integrating the nanometer-level signal interpolation technology of the field programmable gate array with the dynamic resonance suppression algorithm of the digital signal processor, the system positioning accuracy reaches ±7.2 nanometers, achieving positioning accuracy that meets the requirements of 3-nanometer cutting-edge technology. Application measurements show that its worktable positioning error is compressed from 100 nanometers to 10 nanometers, meeting the ±10 nanometer tolerance requirement of the current most advanced process technology. The digital signal processor layer uses a single instruction multiple data instruction set to achieve parallel solution of 16-axis control quantities. The time consumption of model predictive control calculation is greatly reduced compared with the time consumption of traditional solutions, and the efficiency is significantly improved, achieving a breakthrough in the bottleneck of real-time computing efficiency. Combined with the Peripheral Component Interconnect Standard (PCI) Streamlined Frame Protocol, system stability within the specified time control cycle is significantly improved, providing algorithmic support for high-speed nanopositioning. The time-space separation architecture significantly reduces hardware costs compared to dedicated motion controllers through functional integration, significantly shrinking printed circuit board area and significantly lowering power consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0015] Figure 1 is a system architecture diagram provided by an embodiment of the present invention; Figure 2 This is a clock taming network topology diagram provided by an embodiment of the present invention; Figure 3 This is a block diagram of a multi-stage closed-loop control provided by an embodiment of the present invention; Figure 4 It is a schematic diagram of the security protection mechanism provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0016] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are considered to be within the scope of the present invention. Specific implementation method 1 In the time-space separation heterogeneous computing architecture constructed by this invention, the time-sensitive layer, constructed using a field-programmable gate array (FPGA), uses hardware logic to synchronize the rising edge of the global clock for multi-axis analog-to-digital converter sampling, ensuring that the phase deviation of 24-axis sensor data acquisition meets requirements. Simultaneously, it drives the multi-axis pulse-width modulation signals to be strictly aligned and updated in the next clock cycle, significantly reducing the output phase difference. Its emergency stop protection circuit is directly connected to the actuator enable terminal, compressing the response time to 100 nanoseconds, forming a hardware-level safety barrier. The computationally intensive layer, constructed using the AM5728 processor, consists of dual digital signal processor cores and uses a single-instruction, multiple-data instruction set to parallelly process 16-axis control variables, compressing the model predictive control calculation time to 38 microseconds. Combining online mechanical resonance frequency identification with anti-phase sinusoidal compensation technology, it effectively suppresses mechanical resonance, achieving significant algorithm bandwidth.

[0018] The present invention also provides a multi-axis synchronous control method of a multi-axis nanoscale real-time synchronous control system based on heterogeneous computing collaboration, which is characterized by comprising the following steps: Step S1: Configure the OCXO master clock to trigger full system synchronization; Step S2: Use vectorized instructions to calculate the 16-axis control quantities in parallel:

[0019] Where U is the axis control output, E is the axis error vector, K p is the proportional gain matrix, K i is the integral gain matrix, K d is the differential gain matrix; Step S3: The calculation result is transmitted back to the FPGA and all axis output signals are synchronously updated on the rising edge of the next clock cycle. Figure 1 For the system architecture diagram, Figure 1As shown, the oven-controlled crystal oscillator (OCXO) master clock is distributed to 24 ADCs via the LVDS network. The rising edge of the global clock triggers synchronous sampling, with a wiring length error of ≤1mm (delay difference ≤5ps), ensuring that the data acquisition phase deviation is ≤200ps and embeds a ±1ns precision timestamp. The 16-axis PWM output implements parallel loading of the duty cycle through a shared register group and a global reset signal, synchronously updating the control signal on the next rising edge of the clock. A 50Ω impedance matching circuit is deployed at the output end to suppress jitter, and the output phase difference is ≤10ns. The emergency stop protection adopts a physical direct connection mechanism, with the signal directly reaching the FPGA hardware interrupt pin. Within 80ns after triggering, the PWM enable end is forced to close and the mechanical brake is activated.

[0020] The AM5728 dual-DSP cores, based on the Single Instruction Multiple Data instruction set, complete four 32-bit floating-point multiplication and addition operations in a single cycle. The TI-RTOS system divides the dual-core tasks: DSP1 performs position loop PID operations, while DSP2 handles model predictive control optimization. Data is exchanged through shared memory, and a watchdog timer monitors task cycle timeouts and switches to a backup algorithm. Mechanical resonance suppression uses real-time analysis of encoder signals at a 1MHz sampling rate, extracts the resonant main frequency through FFT, and generates an anti-phase compensation signal when the resonant energy exceeds the threshold. Add to PID output and increase differential gain K d ,in To compensate for the signal amplitude, f res is the main frequency of mechanical resonance, t is the precise time reference, and φ is the phase correction; the model predictive control converts the 16-axis state space model into a block diagonal sparse matrix, and accelerates the inversion based on hot start iteration and hardware divider, greatly reducing the calculation time.

[0021] The PCIe protocol uses a streamlined frame structure with a 12-byte header and a 256-byte payload. The FPGA directly writes into the DSP memory mapping area to achieve zero-copy transmission. Dynamic delay compensation is achieved through loopback testing. The FPGA sends a test frame and records time T1. The DSP records time T2 after the frame is sent back. The transmission delay is calculated and the next cycle data transmission time is dynamically adjusted to achieve full system timing alignment. Ultimately, a closed loop of "synchronous sampling → real-time calculation → precise output" is achieved within a 50μs cycle, reducing multi-axis control synchronization errors and improving positioning accuracy.

[0022] System synchronization is based on a clock taming network. The FPGA's integrated oven-controlled crystal oscillator is precisely driven in three paths through a low-voltage differential signal network. The first path is fed to the FPGA's global clock, which uniformly schedules analog-to-digital conversion and pulse-width modulation timing; the second path is connected to the AM5728 phase-locked loop input to achieve timing lock between the processor core and the execution layer; the third path is directly connected to the encoder sampling clock to control the signal acquisition deviation within a predetermined range. The FPGA embeds a high-precision timestamp for each frame of data, and combined with the loopback delay dynamic compensation technology, the timing jitter of the entire system is significantly reduced. Figure 2 As shown, the constant temperature controlled crystal oscillator outputs a 100MHz baseband signal, which is precisely driven in three ways through a low voltage differential signal (LVDS) network. The first way is connected to the FPGA global clock pin through a clock buffer, and serpentine wiring controls the PCB trace length to uniformly schedule the ADC / PWM timing. The second way is input into the AM5728 processor phase-locked loop, configured in integer-N division mode to generate a 3GHz core clock, and dynamic phase detection is used to achieve timing lock between the processor and the execution layer. The third way is directly connected to the optical encoder sampling clock end, transmitted via a twisted shielded pair cable and deployed with an RC filtering network to ensure that the encoder acquisition deviation meets actual requirements.

[0023] The FPGA integrates a time-to-digital converter module, with an OCXO clock driving a 32-bit counter and fine interpolation circuit. It captures the instant ADC sampling is completed and generates a high-precision timestamp, which is written to a fixed location in the data frame header. It automatically performs temperature-voltage compensation every 24 hours, correcting nonlinear errors based on a lookup table to reduce accuracy drift.

[0024] FPGA sends a 128-byte loopback test frame at the start of the control cycle, including the local timestamp T1. DSP returns the original path within the specified time. When FPGA receives the return frame, it records T2 and calculates the value according to the formula Calculate the one-way transmission delay, where t proc Fixed processing delay for DSP; dynamically adjust the data sending time of the next cycle: ,in It is the predicted value of DSP calculation time, and T0 is the benchmark time.

[0025] The LVDS driver deploys spread spectrum modulation to compress the clock jitter caused by power supply noise to within the rated range. The FPGA's built-in eye diagram monitoring module measures the deviation between the ADC sampling edge and the clock rising edge in real time, and feeds this back to the OCXO control DAC for dynamic calibration. Combined with three-way clock coordination and a closed-loop delay compensation, this reduces the timing jitter of the entire system and maintains a small multi-axis signal synchronization error over a wide temperature range.

[0026] like Figure 3As shown in the figure, signal processing and algorithm control jointly ensure accuracy. At the signal processing level, the FPGA performs nanometer-level interpolation on the encoder's original signal, increasing the resolution from 1μm to 0.1nm, forming a position loop, and monitoring the axis offset in real time through a hardware position comparator. At the algorithm control level, the digital signal processor dynamically identifies the mechanical resonant frequency and injects an anti-phase compensation signal to form a speed loop. Combined with the adaptive proportional-integral-differential control algorithm, the gain parameters are adjusted in real time according to changes in load inertia, ultimately achieving stable high-precision multi-axis positioning and forming a current loop.

[0027] The FPGA implements nanometer-level signal interpolation, employing dual-channel quadrature encoder signals to generate high-resolution position data through a sine-cosine lookup table and a CORDIC algorithm. The original 1μm periodic signal is subdivided by a factor of 4096, and combined with noise shaping technology, the effective resolution is increased to 0.1nm. A hardware position comparator monitors axis deviation in real time: a built-in 32-bit subtractor operating at 400MHz compares the difference between the set position and the feedback position. If the difference exceeds the limit, an emergency stop signal is triggered within 50ns. Accuracy is guaranteed by a temperature compensation circuit.

[0028] The DSP performs triple dynamic optimization to achieve precise control, captures encoder vibration signals in real time at a 1MHz sampling rate, and extracts the dominant mechanical resonant frequency through 1024-point FFT analysis. , when the energy in a specific frequency band is greater than 10nm, an anti-phase compensation signal is generated; the load inertia J is estimated synchronously and in real time, where , τ is the motor torque, α is the acceleration, to achieve adaptive PID control, when J changes > 10% proportionally adjust the PID gain: K p Updated to , K i Updated to , K d Updated to , where J new is the current load inertia, J old To update the load inertia before, a resonance suppression-inertia adaptation joint control strategy is formed.

[0029] Signal processing and algorithm control work together to build a nanometer-level precision closed loop. The 0.1nm interpolation data of the FPGA is uploaded to the DSP every 50μs as the PID algorithm error vector E; the inverted compensation signal Superimposed on PID output U in real time.

[0030] The Peripheral Component Interconnect standard protocol utilizes an innovative frame structure with a streamlined 12-byte header and a 256-byte axis data payload, significantly increasing the payload rate and significantly reducing the time required for 16-axis data transmission. Safety protection is implemented collaboratively through the FPGA hardware layer and the digital signal processor software layer. The hardware layer implements a direct connection between the emergency stop signal and the pulse width modulation enable terminal and real-time comparison of out-of-limit positions; the software layer runs a watchdog timer to monitor the control cycle and provides safety status feedback to the FPGA at a 100μs cycle. This creates a complete protection system from the physical layer to the logical layer, with communication optimization and safety mechanisms forming a closed-loop guarantee.

[0031] The PCIe protocol uses an innovative frame structure with a 12-byte streamlined header and a 256-byte payload (16 bytes per axis for storage location / speed / status). The 12-byte streamlined header includes a 4-byte timestamp, a 2-byte axis ID mask, a 2-byte CRC, and a 4-byte control instruction. The 256-byte payload includes 16 bytes per axis for storage location, speed, and status, significantly improving the effective payload rate. Direct write and DMA burst transfers to the DSP memory mapping area are implemented based on the PCIe Gen2 x4 interface. Combined with a packet prefetch mechanism, the 16-axis data transmission time is compressed to 1.5μs. Reed-Solomon error correction code is deployed in the header to significantly reduce the bit error rate.

[0032] like Figure 4 As shown, the emergency stop signal is directly transmitted to the FPGA's dedicated hardware interrupt pin. Once triggered, the PWM enable terminal is quickly pulled low and the motor power supply is cut off. The MOSFET metal oxide semiconductor field effect transistor is turned off in less than 30ns, and the total response time is less than 100ns. The 400MHz hardware position comparator compares the encoder feedback position with the preset safety threshold in real time. If the limit is exceeded, the PWM output is frozen and the brake relay is synchronously activated to form physical-level protection.

[0033] An independent hardware watchdog timer monitors the control cycle and automatically switches to a streamlined PID backup algorithm upon timeout. A 32-bit safety word is generated every 100 μs, with bytes 0-7 indicating the task cycle status; bytes 8-15 containing the memory CRC32 checksum; bytes 16-23 indicating the overload / overtemperature alarm; and bytes 24-31 reserved for transmission to the FPGA via an optimized PCIe frame. A safety shutdown sequence is triggered immediately upon a checksum failure.

[0034] The FPGA cross-verifies the hardware position monitoring results with the position estimate in the DSP safety word. The position estimate is generated by a linear extrapolation algorithm. If the deviation is too large, a second-level emergency stop is initiated. At the moment the emergency event is triggered, the last 128 frames of communication data, including a 64-bit timestamp and axis status, are frozen and written into the ferroelectric memory to build a complete fault traceability chain. Specific implementation method 2 The core hardware of the four-axis control system for the lithography machine's worktable is based on an AM5728 processor and a field-programmable gate array (FPGA). The AM5728 processor, with dual digital signal processor (DSP) cores operating at 750 MHz, is responsible for executing the core control algorithms. The FPGA handles high-speed, real-time signal acquisition, synchronization, and output drive tasks. The system uses an ultra-precise clock source to provide a global clock signal with a frequency stability of ±25 parts per billion (ppb), ensuring a high-precision timing benchmark for the entire system. High-speed data exchange between the processor and FPGA is achieved via a Peripheral Component Interconnect Express (PCIe) Gen2x4 interface, which offers a transfer rate of up to 5 GT / s, ensuring low-latency transmission of control commands and feedback data along the critical path.

[0036] The control process strictly adheres to a fixed 50-microsecond cycle. Each control cycle begins with the rising edge of the field-programmable gate array (FPGA) global clock, triggering the synchronous sampling of position and velocity signals for the four motion axes (X, Y, Z, and Rz). The sampled data is then transmitted to the digital signal processor (DSP) via a high-speed PCIe channel. This data transfer is completed within approximately 1.2 microseconds of sampling initiation. Upon receiving the data, the DSP immediately executes a complex model predictive control (MPC) algorithm, which takes approximately 22 microseconds to complete. The resulting optimal control instructions are transmitted back to the FPGA 23 microseconds after the cycle begins. The FPGA receives and processes these instructions, preparing to update the control signals for each axis. Finally, at the next rising edge of the global clock (i.e., 50 microseconds after the cycle begins), the FPGA synchronously updates the pulse-width modulated (PWM) drive signals output to the four axes, thereby driving the workpiece stage actuators for precise motion.

[0037] After rigorous testing, the four-axis control system demonstrated exceptional performance. In terms of synchronization performance, the synchronization error of each axis was kept to extremely low levels: 7.2 nanoseconds for the X-axis, 8.1 nanoseconds for the Y-axis, 5.8 nanoseconds for the Z-axis, and 6.3 nanoseconds for the Rz-axis. Positioning accuracy also reached the nanometer level: ±6.5 nanometers for the X-axis, ±7.1 nanometers for the Y-axis, ±4.9 nanometers for the Z-axis, and ±5.7 nanometers for the Rz-axis. These measured data fully validate the effectiveness of the system's hardware design, control algorithms, and timing management.

[0038] The system's core control cycle is set at 50 microseconds, providing a foundation for balancing real-time performance and control accuracy. The pulse-width modulation (PWM) output, driven by a 400 MHz high-frequency clock, achieves a resolution of up to 2.5 nanoseconds, ensuring precise position control at the micron and even nanometer levels. Regarding safety mechanisms, the system incorporates a hardware-level emergency stop response link. When an emergency stop signal is triggered, the system forcibly shuts off the pulse-width modulation (PWM) outputs of all axes in an extremely short time (less than 80 nanoseconds), ensuring an immediate stop of worktable motion and maximizing equipment and operator safety.

[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A multi-axis nanoscale real-time synchronous control system based on heterogeneous computing collaboration, characterized in that: The system includes the following modules: A time-space separation architecture module, comprising a computationally intensive layer comprised of dual digital signal processor cores of the AM5728 processor and a time-sensitive layer comprised of an FPGA; A global synchronization mechanism module, the global synchronization mechanism module includes an oven-controlled crystal oscillator clock source, the oven-controlled crystal oscillator clock source signal is distributed to the FPGA, AM5728 and encoder via a low-voltage differential signal network; Collaborative workflow module, which implements the following process: (a) FPGA synchronously collects multi-axis data and adds time stamps on the rising edge of the clock; (b) Transmitting the multi-axis data with time stamps to the DSP via the PCIe or GPMC interface; (c) Dual digital signal processors complete control algorithm calculations within the specified time; (d) The calculation results are transmitted back to the FPGA and the multi-axis outputs are synchronously updated on the next rising edge of the clock.

2. The system according to claim 1, wherein: The FPGA includes: 24-channel ADC synchronous sampling unit, wherein the sampling jitter of the ADC synchronous sampling unit is less than a first threshold; 16-axis PWM synchronous output unit, updated by the same clock edge trigger; An emergency stop protection circuit, wherein a response time of the emergency stop protection circuit is less than a second threshold.

3. The system according to claim 1, wherein: The DSP core runs the TI-RTOS real-time operating system; vectorized instructions are used for parallel computing to achieve computational acceleration; and an integrated watchdog timer monitors whether a task cycle has timed out.

4. The system according to claim 1, wherein: The specific structure of the PCIe or GPMC interface is: The PCIe protocol frame structure consists of a 12-byte header and a 256-byte payload, with a bandwidth utilization rate of 95%; or The GPMC interface timing parameters are address line setup time 4ns and data line hold time 2ns.

5. The system according to claim 1, wherein: The FPGA performs nanometer-level interpolation on the collected multi-axis data and monitors axis offset in real time through hardware position comparators; Dual digital signal processors dynamically identify the mechanical resonant frequency in multi-axis data and inject anti-phase compensation signals. Combined with an adaptive proportional-integral-derivative control algorithm, gain parameters are adjusted in real time according to changes in load inertia.

6. A multi-axis synchronous control method for a multi-axis nanoscale real-time synchronous control system based on heterogeneous computing collaboration as described in claims 1-5, characterized in that: The following steps are involved: Step S1, configuring the oven-controlled crystal oscillator master clock to trigger full system synchronization; Step S2, use vectorized instructions to calculate the 16-axis control quantities in parallel: ; Where U is the axis control output, E is the axis error vector, is the proportional gain matrix, is the integral gain matrix, is the differential gain matrix, t represents time; In step S3, the calculation results are transmitted back to the FPGA and all axis output signals are synchronously updated on the rising edge of the next clock cycle.

Citation Information

Patent Citations

  • Multi-module multi-channel acquisition synchronization system and the working method thereof

    CN106406174A

  • Multi-axis linkage embedded type digital control system and development method thereof

    CN108549330A

  • Multi-axis high-precision space motion platform

    CN117008275A

  • Multi-shaft electronic gear synchronous control system and method based on FPGA

    CN118192381A

  • Synchronous control method, device and equipment for multi-axis servo system

    CN119937330A

Cited By

  • Transverse movement cooperative control system and method for ultra-high-speed warp knitting machine

    CN122219286A

  • Transverse movement cooperative control system and method for super high speed warp knitting machine

    CN122219286B