Synchronization device for GPU and image processing system

By designing a synchronization device for the GPU, the first processing unit and the second processing unit receive and output signals, the signal synchronization of multiple GPUs is realized, and the problem of complex and high cost synchronization of video wall systems in the prior art is solved, and a low-cost and easy-to-scaling synchronization solution is realized.

CN119785690BActive Publication Date: 2025-06-10MOORE THREADS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510273027.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-10
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

Existing video wall systems require dedicated displays and signal control systems to achieve signal synchronization of each display, which is low in versatility, high in cost and complex in deployment.

Method used

A GPU synchronization device is designed, including a first processing unit and a second processing unit. By receiving an effective clock signal and a frame synchronization signal, outputting the frame data to be displayed and the effective synchronization signal is realized, and signal synchronization of multiple GPUs is realized.

Benefits of technology

The signal synchronization of multiple GPUs is achieved without the need for a dedicated display system or data processing system, reducing costs and simplifying deployment, suitable for a variety of video walls and multi-display scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785690B_ABST
    Figure CN119785690B_ABST
Patent Text Reader

Abstract

The present disclosure provides a synchronization device for a GPU and an image processing system. The synchronization device includes: a first processing unit configured to receive a preset valid clock signal and output the valid clock signal to at least one GPU; a second processing unit configured to receive a first frame synchronization signal of the at least one GPU and output a second frame synchronization signal to the at least one GPU; wherein the second frame synchronization signal is a valid synchronization signal when the first frame synchronization signals all indicate that the GPU is in a frame synchronization state; the valid synchronization signal is used to indicate that the at least one GPU is in a frame synchronization state; and the GPU outputs frame data to be displayed based on the valid clock signal when the second frame synchronization signal is the valid synchronization signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and in particular, to a synchronization device for a GPU (Graphics Processing Unit) and an image processing system. Background Art

[0002] A video wall is an image and graphic display system, which is a super-large screen video display wall composed of multiple display units spliced together. Video walls have a wide range of applications in many fields, including but not limited to places such as monitoring centers, command centers, video conference rooms, exhibition halls, etc., for displaying real-time video monitoring, multimedia information, data analysis results, etc. A video wall usually consists of splicing units, multi-screen processors, signal switching and distribution, control systems, etc. Among them, the multi-screen processor is often the core of splicing the video wall, and is used to divide a complete image signal and distribute it to each video display unit.

[0003] However, video walls often need to be composed of dedicated displays, and at the same time require the cooperation of the sending end, or use a dedicated signal control system to achieve signal synchronization of each display, with low versatility, high cost and complex deployment. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide a synchronization device for a GPU and an image processing system to solve at least one problem existing in the prior art.

[0005] To achieve the above object, the technical solution of the embodiments of the present disclosure is implemented as follows: On the one hand, embodiments of the present disclosure provide a synchronization device for a GPU, including:

[0006] A first processing unit, configured to receive a preset valid clock signal and output the valid clock signal to at least one GPU;

[0007] A second processing unit, configured to receive a first frame synchronization signal of the at least one GPU and output a second frame synchronization signal to the at least one GPU; wherein, the second frame synchronization signal is a valid synchronization signal when the first frame synchronization signals all indicate that the GPU is in a frame synchronization state; the valid synchronization signal is used to indicate that the at least one GPU is in a frame synchronization state;

[0008] The GPU outputs frame data to be displayed based on the valid clock signal when the second frame synchronization signal is the valid synchronization signal.

[0009] In some embodiments, the first processing unit includes:

[0010] An OR operation unit includes multiple input terminals and is used to perform OR operation processing on the signals received by the multiple input terminals;

[0011] At least some of the multiple input terminals of the OR operation unit are respectively used to connect to the GPU;

[0012] Wherein, any one of the GPUs inputs the valid clock signal to the OR operation unit, or a preset external input inputs the valid clock signal to the OR operation unit through the input terminal;

[0013] The output terminal of the OR operation unit is connected to the at least one GPU and is used to provide the valid clock signal to the at least one GPU.

[0014] In some embodiments, the multiple input terminals of the OR operation unit are respectively connected to pull-down resistors; the pull-down resistors are used to set the input signal of the input terminal to logic 0 when the valid clock signal is not received.

[0015] In some embodiments, the first processing unit includes:

[0016] A signal selection unit includes multiple input terminals and is used to select the valid clock signal among the signals received by the multiple input terminals as the output signal;

[0017] At least some of the multiple input terminals of the signal selection unit are respectively used to connect to the GPU;

[0018] Wherein, any one of the GPUs inputs the valid clock signal to the signal selection unit, or a preset external input inputs the valid clock signal to the signal selection unit through the input terminal;

[0019] The output terminal of the signal selection unit is connected to the at least one GPU and is used to provide the valid clock signal to the at least one GPU.

[0020] In some embodiments, the second processing unit includes:

[0021] An AND operation unit;

[0022] The AND operation unit includes multiple input terminals; at least some of the multiple input terminals of the AND operation unit are respectively used to connect to the GPU and receive the first frame synchronization signal provided by the GPU;

[0023] Wherein, the high-level first frame synchronization signal is the valid synchronization signal.

[0024] In some embodiments, the multiple input terminals of the AND operation unit are respectively connected to pull-up resistors; the pull-up resistors are used to set the input terminal not connected to the GPU to logic 1.

[0025] In some embodiments, the second processing unit includes:

[0026] A counting unit;

[0027] The counting unit includes a plurality of input terminals; at least some of the plurality of input terminals of the counting unit are respectively used to connect to the GPU and receive the first frame synchronization signal provided by the GPU;

[0028] The counting unit is configured to count the valid synchronization signals in the received first frame synchronization signal, and output the second frame synchronization signal to the at least one GPU when the counted value is consistent with the number of received first frame synchronization signals.

[0029] In some embodiments, the synchronization device includes N GPU interfaces; the GPU interfaces at least include N first connection terminals connected to the first processing unit and N second connection terminals connected to the second processing unit; wherein, N is an integer greater than or equal to 2;

[0030] The synchronization device is connected to M of the GPUs through the GPU interfaces, and is configured to provide the valid clock signal to the M GPUs through the first processing unit, and provide the second frame synchronization signal to the M GPUs through the second processing unit;

[0031] wherein, M is a positive integer less than or equal to N.

[0032] In some embodiments, the synchronization device further includes:

[0033] A cascade input terminal and a cascade output terminal;

[0034] The cascade input terminal is connected to the cascade output terminal of another cascaded synchronization device;

[0035] The cascade output terminal and the cascade input terminal are configured to transmit the valid clock signal and / or the valid synchronization signal between the cascaded plurality of synchronization devices.

[0036] In some embodiments, the cascade input terminal includes: a clock cascade input terminal; the cascade output terminal includes: a clock cascade output terminal;

[0037] The clock cascade input terminal is connected to the input terminal of the first processing unit and is configured to connect to the clock cascade output terminal of another cascaded synchronization device;

[0038] The clock cascade output terminal is connected to the output terminal of the first processing unit.

[0039] In some embodiments, when the clock cascade input end is not connected to the clock cascade output end of the other synchronization device, the clock cascade input end floats; and / or

[0040] When the clock cascade output end floats and is not connected to the clock cascade input end of the other synchronization device, the clock cascade output end floats.

[0041] In some embodiments, the cascade input end includes: a first synchronization input end; the cascade output end includes: a first synchronization output end;

[0042] The first synchronization input end is connected to the input end of the second processing unit and is used to connect the first synchronization output end of another synchronization device in cascade;

[0043] The first synchronization output end is connected to the output end of the second processing unit.

[0044] In some embodiments, the cascade input end further includes: a second synchronization input end, and the cascade output end further includes: a second synchronization output end;

[0045] Both the second synchronization input end and the second synchronization output end are connected to the at least one GPU for transmitting the valid synchronization signal, and the second synchronization input end is used to connect the second synchronization output end of another synchronization device in cascade, or is used to connect the first synchronization output end of the synchronization device.

[0046] In some embodiments, when no other synchronization device is cascaded on the first side of the synchronization device and another synchronization device is cascaded on the second side:

[0047] Both the first synchronization input end and the second synchronization output end float;

[0048] The first synchronization output end is connected to the first synchronization input end of the other synchronization device;

[0049] The second synchronization input end is connected to the second synchronization output end of the other synchronization device.

[0050] In some embodiments, when a first synchronization device is cascaded on the first side of the synchronization device and a second synchronization device is cascaded on the second side:

[0051] The first synchronization input end is connected to the first synchronization output end of the first synchronization device;

[0052] The second synchronization output end is connected to the second synchronization input end of the first synchronization device;

[0053] The first synchronization output end is connected to the first synchronization input end of the second synchronization device;

[0054] The second synchronization input terminal is connected to the second synchronization output terminal of the second synchronization device;

[0055] Wherein, the first synchronization device and the second synchronization device are the same as the synchronization device.

[0056] In some embodiments, when another synchronization device is cascaded on the first side of the synchronization device and no other synchronization device is cascaded on the second side:

[0057] The first synchronization input terminal is connected to the first synchronization output terminal of the another synchronization device;

[0058] The second synchronization output terminal is connected to the second synchronization input terminal of the another synchronization device;

[0059] The first synchronization output terminal is connected to the second synchronization input terminal.

[0060] In some embodiments, when no other synchronization device is cascaded on the synchronization device:

[0061] Both the first synchronization input terminal and the second synchronization output terminal are floating;

[0062] The first synchronization output terminal is connected to the second synchronization input terminal.

[0063] In some embodiments, the synchronization device further includes:

[0064] One or more synchronization indication units; the synchronization indication unit is configured to receive the clock synchronization status signal sent by the GPU based on the valid clock signal, and indicate that the GPU is in the clock synchronization state.

[0065] On the other hand, an embodiment of the present disclosure provides an image processing system, including:

[0066] One or more cascaded synchronization devices of any of the above;

[0067] Multiple GPUs, and each synchronization device is connected to one or more of the GPUs.

[0068] In some embodiments, the image processing system further includes:

[0069] A display module, connected to the multiple GPUs;

[0070] The GPU outputs the frame data to be displayed to the display module based on the valid clock signal and the valid synchronization signal output by the synchronization device;

[0071] The display module displays the spliced picture based on the frame data output by the multiple GPUs.

[0072] In the technical solution provided by the present disclosure, a synchronization device for a GPU is provided. The device processes the received valid clock signal through a simple first processing unit and a second processing unit and provides it to at least one GPU so that at least one GPU receives the same valid clock signal as the reference clock for output image data; and provides a second frame synchronization signal to each GPU when a valid synchronization signal indicating the frame synchronization state is sent by one or more connected GPUs to inform each GPU of the current synchronization state, thereby realizing the display synchronization of each GPU. In this way, the synchronization device can conveniently realize the signal synchronization of multiple GPUs without a dedicated display system or data processing system, with low cost and easy expansion, and can be applied to various video walls, scenarios requiring multi-display display, and scenarios requiring image processing by multiple GPUs. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 Structural schematic of a synchronization device for a GPU provided by an embodiment of the present disclosure Figure 1 ;

[0074] Figure 2 Structural schematic of a synchronization device for a GPU provided by an embodiment of the present disclosure Figure 2

[0075] Figure 3 Structural schematic of a synchronization device for a GPU provided by an embodiment of the present disclosure Figure 3 ;

[0076] Figure 4 Structural schematic of a synchronization device for a GPU provided by an embodiment of the present disclosure Figure 4 ;

[0077] Figure 5 Schematic diagram of the connection method for the synchronization device provided by an embodiment of the present disclosure to achieve cascading;

[0078] Figure 6 Schematic diagram of the connection method for the synchronization device provided by an embodiment of the present disclosure without cascading;

[0079] Figure 7 Schematic diagram of the wiring of the synchronization device provided by an embodiment of the present disclosure applied in a system architecture;

[0080] Figure 8 Flowchart of the configuration method of the synchronization device and multiple GPUs provided by an embodiment of the present disclosure;

[0081] Figure 9 Structural schematic of the synchronization device provided by an embodiment of the present disclosure applied to a display system;

[0082] Figure 10 Block diagram of an image processing system provided by an embodiment of the present disclosure;

[0083] Figure 11 Block diagram of a GPU provided by an embodiment of the present disclosure. Detailed implementation manners

[0084] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the specific embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0085] In the following description, numerous specific details are given to provide a more thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure can be practiced without one or more of these details. In other instances, some well-known technical features are not described in order to avoid obscuring the present disclosure; that is, not all features of the actual embodiments are described here, and the well-known functions and structures are not described in detail.

[0086] The purpose of the terms used herein is only to describe specific embodiments and is not a limitation of the present disclosure. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprising" and / or "including", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups. As used herein, the term "and / or" includes any and all combinations of the related listed items.

[0087] To achieve synchronous display of a multi-screen video wall, an independent video wall system (implemented in hardware or software) can be used to collect, cut, and send video to multiple displays. The display end is required to support the frame synchronization signal sent by the video wall system, thereby achieving synchronous display of multiple displays. This method requires an additional dedicated system or an additional dedicated synchronization signal for the display, which lacks generality and increases the usage cost and deployment difficulty. Alternatively, a synchronization signal can be added at the display end to ensure that the received pictures can be synchronously displayed. However, this method requires a dedicated display and also requires cooperation from the sending end during actual use to avoid excessive synchronization delays among the displays.

[0088] As Figure 1 shown, an embodiment of the present disclosure provides a synchronization device 100 for a GPU, including:

[0089] The first processing unit 110 is configured to receive a preset valid clock signal clk and output the valid clock signal clk to at least one GPU.

[0090] The second processing unit 120 is configured to receive the first frame synchronization signal fsync1 of the at least one GPU and output a second frame synchronization signal fsync2 to the at least one GPU; wherein, the second frame synchronization signal fsync2 is a valid synchronization signal when the first frame synchronization signals fsync1 all indicate that the GPU is in a frame synchronization state; the valid synchronization signal is used to indicate that the at least one GPU is in a frame synchronization state.

[0091] The GPU outputs frame data to be displayed based on the valid clock signal clk when the second frame synchronization signal fsync2 is the valid synchronization signal.

[0092] The above-mentioned first processing unit 110 may have multiple signal channels for receiving clocks, and receive one or more clock signals through these signal channels, and select the valid clock signal clk as the output of the first processing unit 110 to be provided to each GPU connected to the synchronization device 100.

[0093] The valid clock signal may come from any one of the GPUs connected to the synchronization device 100. For example, one of the GPUs can be preset as the main GPU through the processor or other devices connected to each GPU to output the valid clock signal, while the other GPUs do not output clock signals or output invalid signals.

[0094] The valid clock signal may also come from other signal sources connected to the synchronization device 100. For example, a fixed clock signal output by a preset clock signal generator is directly input to the synchronization device 100 as the valid clock signal; or the clock signal provided by a processor externally connected to the synchronization device is used as the valid clock signal.

[0095] The valid clock signal may also come from other synchronization devices cascaded with the synchronization device 100, that is, the same valid clock signal can be used after multiple synchronization devices are cascaded. The specific cascading method can refer to the specific description in the relevant embodiments hereinafter.

[0096] It should be noted that for each GPU, a fixed effective clock signal can be used throughout the entire process after power-on. That is, after each power-on or initialization operation, the effective clock signal is synchronized once, and then regardless of whether a video frame is transmitted, the synchronization of the effective clock signal can be maintained. For the transmission of video frames, each GPU needs to synchronize each frame separately. When the GPU has completed the processing of a video frame, such as decoding and rendering, and can send the processed data to the display module for display, the GPU can output a corresponding frame synchronization signal to indicate that the current GPU has completed the processing of the current frame. Therefore, in the embodiments of the present disclosure, the synchronization device 100 can use the frame synchronization signal sent by the GPU to determine whether multiple GPUs have all completed the processing of the current frame, and after multiple GPUs have completed the current processing, instruct each GPU to output the data of the current frame for the display device to display.

[0097] The second processing unit 120 is used to determine whether each GPU connected to the synchronization device 100 is in a synchronized state. The second processing unit 120 receives the first frame synchronization signals fsync1 output by each connected GPU, and this signal is used to indicate whether the corresponding GPU has completed the processing of the current frame. The second processing unit 120 can be connected to one or more GPUs and receive the first frame synchronization signals fsync1 output by each of the connected GPUs, and output a second frame synchronization signal to each GPU. It can be understood that in order to achieve the synchronization of the output of the same video frame by each GPU, when the first frame synchronization signals fsync1 output by each GPU all indicate that the GPU is in a frame synchronization state (or understood as having completed processing and being able to send data), the output second frame synchronization signal is a valid synchronization signal. For example, if the first frame synchronization signal is at a high level, it indicates that the GPU is in a frame synchronization state. When the first frame synchronization signals output by each connected GPU are all at a high level, the second processing unit 120 outputs a valid synchronization signal (that is, a valid second frame synchronization signal, such as a high level). Correspondingly, if the first frame synchronization signal output by any at least one connected GPU is at a low level, indicating that the GPU has not completed processing and is temporarily unable to send data, the second processing unit 120 outputs a second frame synchronization state indicating invalidity (that is, an unsynchronized state). The above "high level" or "low level" are only exemplary embodiments. In practical applications, the manifestation forms of the above signals are not limited to this form, and can also represent whether the GPU is in a valid state of completing the processing of a frame of data and waiting to send data through different signal waveforms, signal duty cycles, signal frequencies, and signal on / off. The output signal of the second processing unit 120 can also adopt any of the above manifestation forms of signals, which will not be elaborated here.

[0098] In some embodiments, the synchronization device 100 includes N GPU interfaces; the GPU interfaces at least include N first connection ends 111 connected to the first processing unit 110 and N second connection ends 121 connected to the second processing unit 120; wherein, N is an integer greater than or equal to 2;

[0099] The synchronization device is connected to M of the GPUs through the GPU interfaces and is configured to provide the valid clock signal to the M GPUs through the first processing unit, and provide the second frame synchronization signal to the M GPUs through the second processing unit;

[0100] wherein, M is a positive integer less than or equal to N.

[0101] In the embodiments of the present disclosure, the number of GPU interfaces of the synchronization device 100 can be designed according to actual application requirements. For example, if a GPU synchronization device for conventional household or small devices is provided, such as a device with 1 to 4 GPUs, a synchronization device with 4 GPU interfaces can be designed. For application scenarios such as large video walls or AI (Artificial Intelligence) model training devices (or device clusters) that require high computing power, for example, in the case where 32 or 64 GPUs are required to process an image processing task simultaneously, the synchronization device can be designed with 32 or 64 GPU interfaces, or relatively fewer GPU interfaces, such as 8. During use, cascading can be performed according to requirements to expand the number of connectable GPUs and achieve flexible multi-GPU synchronization configuration.

[0102] Of course, in some scenarios, the number of GPUs required to be used may be less than the number of GPU interfaces of the synchronization device. For example, a synchronization device with 4 GPU interfaces can be connected to only 1 GPU for use, or can be connected to 2 to 4 GPUs for use and achieve the synchronization of these 2 to 4 GPUs.

[0103] In some embodiments, as Figure 2 shown, the above-mentioned first processing unit 110 includes: an OR operation unit 110a; the OR operation unit 110a is configured to perform OR operation processing on signals received by multiple input ends, and at least some of the multiple input ends of the OR operation unit 110a are respectively used to connect to GPUs. Exemplarily, the multiple input ends of the OR operation unit 110a are respectively connected to multiple GPUs (such as Figure 2The reference clock output terminals of GPU0 to GPU3 shown in the figure; the output terminal of the OR operation unit 110a is connected to the reference clock input terminal of each of the GPUs; wherein, at most one of the multiple reference clock output terminals is used to provide a reference clock output signal o_refclk; the reference clock input terminal is used to provide a valid clock signal i_refclk to the GPU.

[0104] In some embodiments, multiple input terminals of the above OR operation unit 110a can be respectively connected to pull-down resistors. The pull-down resistors are used to set the input signal of the connected input terminal to logic "0" when no valid clock signal is received. In this way, even if the input terminal is not connected to a GPU, or the connected GPU does not output any signal, the OR operation unit 110a can still operate normally. It can be understood that if the OR operation unit 110a is composed of one or more OR gates, each input terminal of the OR gate needs to provide an input signal and cannot be floating, otherwise the accuracy of the output result may be affected due to interference on the floating input terminal. Therefore, in the embodiments of the present disclosure, each input terminal is set to logic "0" by a pull-down resistor, that is, when there is no signal input, the input terminal defaults to input logic "0". When a valid signal logic "1" or logic "0" is input, it will be input through the input terminal, and the pull-down resistor will be short-circuited.

[0105] In addition, in some embodiments, as Figure 2 shown, the above second processing unit 120 includes: an AND operation unit 120a. The AND operation unit 120a also includes multiple input terminals, and at least some of the input terminals are respectively used to connect to a GPU and receive a first frame synchronization signal provided by the GPU. Exemplarily, the multiple input terminals of the AND operation unit 120a are respectively connected to the synchronization signal output terminals of the multiple GPUs; the output terminal of the AND operation unit 120a is connected to the synchronization signal input terminal of each of the GPUs; wherein, the synchronization signal output terminal is used to output a frame synchronization output signal o_fsync of the GPU; the synchronization signal input terminal is used to provide a valid synchronization signal i_fsync to the GPU;

[0106] wherein, the GPU outputs frame data to be displayed based on the valid clock signal i_refclk and the valid synchronization signal i_fsync.

[0107] In some embodiments, a plurality of input terminals of the above-mentioned AND operation unit 120a may be respectively connected to pull-up resistors; the pull-up resistors are used to set the input terminals not connected to the GPU to logic "1". In this way, even if the input terminal is not connected to the GPU, or the connected GPU does not output any signal, the AND operation unit 120a can still operate normally. It can be understood that if the AND operation unit 120a is composed of one or more AND gates, then similar to the above-mentioned OR gates, each input terminal of the AND gate needs to provide an input signal and cannot be floating, otherwise the accuracy of the output result may be affected due to interference on the floating input terminal. Therefore, in the embodiments of the present disclosure, each input terminal is set to logic "1" through a pull-up resistor, that is, in the case of no signal input, the input terminal defaults to input logic "1". When there is a logic "1" or logic "0" input from the GPU, it will be input to the AND operation unit 120a through the input terminal, and the pull-up resistor will be short-circuited.

[0108] For the above Figure 2 In the case shown, the OR operation unit performs an "OR" logic operation on the input signals received by the plurality of input terminals and provides an output signal. Among the input signals of the plurality of input terminals, if at least one is logic "1", that is, high level, then the output is logic "1", that is, high level. If the input signals of the plurality of input terminals are all logic "0", that is, low level, then the output is logic "0", that is, low level.

[0109] It should be noted that in the embodiments of the present disclosure, the reference clock input terminal of the GPU is used to receive the valid clock signal i_refclk, which is the reference clock signal used by the GPU. When the GPU performs image processing, it can output image data according to the clock transition edge of the valid clock signal i_refclk and provide it to the corresponding display device. Therefore, in order to synchronize the clocks of multiple GPUs, a synchronized valid clock signal i_refclk needs to be provided to the reference clock input terminals of multiple GPUs. And each GPU itself can output the internal clock signal. Therefore, one GPU can be selected as the main card to provide the reference clock signal.

[0110] In the embodiments of the present disclosure, the plurality of input terminals of the OR operation unit 110a are used to receive the reference clock signals provided by the reference clock output terminals of multiple GPUs. In order to synchronize the clocks of multiple GPUs, one of the GPUs is selected as the main card, and its output reference clock output signal is provided to the OR operation unit 110a as one of the input signals. The other GPUs are used as slave cards, and the reference clock output terminals of the slave cards can output low-level signals, that is, logic "0", or do not output any signal.

[0111] In this way, the output terminal of the OR operation unit 110a can output the reference clock output signal o_refclk output from the reference clock output terminal of the main card as the reference clock signal, and is connected to the output terminal of the OR operation unit 110a through the reference clock input terminals of multiple GPUs to synchronously receive the reference clock signal, that is, the above-mentioned valid clock signal i_refclk. In this way, the clock synchronization of multiple GPUs can be achieved through a simple OR operation unit.

[0112] It can be understood that in order for the OR operation unit 110a to normally output the reference clock signal, only one of its multiple inputs is the clock signal, otherwise it will cause abnormal output. In some embodiments, in addition to the multiple input terminals connected to the reference clock output terminals of the above-mentioned multiple GPUs, the multiple input terminals of the OR operation unit 110a may further include a reference clock input terminal connected to an external input. If there is a main card among the above-mentioned multiple GPUs for providing the reference clock output signal o_refclk, the externally input reference clock input terminal can be disconnected from the card, without any input signal, or grounded or connected to a low-level input signal. If there is no main card among the above-mentioned multiple GPUs for providing the reference clock output signal o_refclk, the reference clock output signal o_refclk can be provided through the externally input reference clock input terminal.

[0113] It can be understood that the reference clock input terminal can be used to connect to the reference clock output terminal of other GPUs, can also be used to connect to a clock signal generator dedicated to outputting clock signals, and can also be used to connect to the output terminal of the OR operation unit 110a of other synchronization devices to achieve the expansion of the synchronization device.

[0114] The above AND operation unit 120a is a module that performs an "AND" logic operation on the input signals received by multiple input terminals and provides an output signal. Among the input signals of multiple input terminals, if at least one is logic "0", that is, low level, then logic "0", that is, low level, is output. Only when all input signals are logic "1", that is, high level, will logic "1", that is, high level, be output.

[0115] Since the GPU needs a certain preparation time to process the data of each frame of the picture, when the data of a frame of the picture is prepared and stored in the buffer, it means that the signal can be output to the display end for display. The above-mentioned synchronization signal output terminal of the GPU is used to output the frame synchronization output signal o_fsync. Therefore, after each frame of the picture is processed and cached, the synchronization signal output terminal can be set to "1", that is, output the high-level frame synchronization output signal o_fsync, indicating that the GPU has completed the data preparation of the current frame.

[0116] When the frame synchronization output signals o_fsync received by multiple input terminals of the AND operation unit 120a of the synchronization device 100 are all at the high level logic "1", it indicates that all GPUs have completed the data preparation for the current frame. At this time, if the synchronization of the above-mentioned valid clock signal i_refclk has been completed, each GPU can be instructed to output frame data to complete the display of the current frame.

[0117] Since the output terminals of the above-mentioned AND operation unit 120a are respectively connected to the synchronization signal input terminals of multiple GPUs, the AND operation result can be used as an indication signal for frame synchronization and provided to each GPU. When a GPU receives the above-mentioned valid synchronization signal i_fsync at the high level, i.e., logic "1", it indicates that multiple GPUs have all completed the preparation of frame data, so that frame data can be output for the display terminal to display.

[0118] It should be noted that the above clock synchronization can be achieved when the device is powered on or during the initialization process of preparing to start video display. When multiple GPUs complete the synchronization of the reference clock (that is, each GPU's reference clock input terminal receives the valid clock signal i_refclk), the multiple GPUs complete the clock synchronization. Thereafter, the clock synchronization can be maintained unchanged to ensure the clock synchronization of subsequent frame displays. The synchronization of the above frame data is separately implemented after each GPU completes the data preparation for each frame, that is, frame synchronization needs to be performed for each frame display. After receiving the valid synchronization signal i_fsync input at the synchronization signal input terminal, the GPU outputs the current frame of data, and then the output signal at the synchronization signal output terminal can be cleared, that is, the frame synchronization output signal o_fsync is set to "0", until the data for the next frame is completed and the frame synchronization output signal o_fsync is reset to "1". In this way, on the premise of clock synchronization, the frame synchronization of each frame can be completed in sequence, and the display of the stitched picture can be completed by outputting each frame in sequence.

[0119] In this way, by using the above synchronization device 100, the clock synchronization and frame synchronization operations of multiple GPUs can be realized by simply connecting multiple GPUs. Moreover, there is no need for communication between each GPU, so there is no need to set up an additional synchronization signal channel, nor is it necessary to configure a dedicated display device. The display terminal can be multiple different display units, and each display unit can be separately connected and controlled with multiple GPUs. For each GPU and the display unit, there is no need to consider synchronization separately. The synchronization device 100 has completed the clock synchronization and frame synchronization of multiple GPUs, and the display terminal can display according to the received data and clock signals. Therefore, the above synchronization device has stronger versatility and can be widely applied to various scenarios such as various video walls, multi-screen displays, and picture segmentation, with low cost and easy expansion.

[0120] In some other embodiments, the above-mentioned first processing unit 110 and second processing unit 120 may also be other types of logic operation units that can achieve similar functions. For example, the above-mentioned first processing unit 110 may also be a signal selection unit. For example, the signal selection unit may be composed of one or more signal selectors. At least some of the multiple input terminals of the signal selection unit are respectively used to connect to the GPU. The signal selection unit selects a valid clock signal from the received multiple signals as the output signal.

[0121] Exemplarily, the signal selection unit may further include a signal selection terminal, which can be used to receive a pre-set selection signal, and the selection signal is used to indicate which input signal the signal selection unit selects as the output signal. For example, when it is configured that the first input terminal receives a valid clock signal, the selection signal can be configured as the signal corresponding to the first input terminal, so that the signal selection unit can output the signal of the first input terminal. When it is modified to that the second input terminal receives a valid clock signal, the selection signal can be configured as the signal corresponding to the second input terminal, so as to switch the output signal to the input signal of the second input terminal.

[0122] For another example, the second processing unit 120 may also be a counting unit, and the counting unit may be composed of one or more counters. The counting unit can count the valid synchronization signals in the received input signals. When the count value is consistent with the number of the received first frame synchronization signals, it indicates that all the received first frame synchronization signals are valid synchronization signals, and thus a second frame synchronization signal can be output to indicate that each GPU is in a frame synchronization state. For example, the synchronization device 100 is connected to 3 GPUs, and each GPU outputs a first frame synchronization signal to the counting unit respectively. When the first frame synchronization signal is invalid, it is at a low level, and when it is valid, it is at a high level. Therefore, the counting unit can count the received high levels. When the count value is 3, a second frame synchronization signal is output. For another example, the synchronization device 100 has a total of 4 GPU interfaces, 3 of which are respectively connected to a GPU, and the input terminal of the counting unit that is not connected to the GPU is set to logic "1" (for example, a high level is input through a pull-up resistor). In this way, the counting unit counts the high levels received at each input terminal. When the count value is consistent with the total number of GPU interfaces (the count value is 4), a second synchronization signal is output.

[0123] In some embodiments, the synchronization device 100 further includes: a cascaded input terminal and a cascaded output terminal; the cascaded input terminal is connected to the cascaded output terminal of another cascaded synchronization device; the cascaded output terminal and the cascaded input terminal are used to transmit the valid clock signal and / or the valid synchronization signal between multiple cascaded synchronization devices.

[0124] It can be understood that the above synchronization device 100 realizes clock synchronization by the first processing unit 110 and frame synchronization by the second processing unit 120. Both of these logic units can be implemented by one or more logic gates. Therefore, different data volume input ends can be implemented through the setting of logic gates, and the cascading of multiple synchronization devices can be achieved through the cascading of the input ends and output ends of the logic units. In this way, doubling expansion can be achieved through cascading, which is convenient for flexibly setting the number of GPUs.

[0125] In some embodiments, as Figure 3 shown, the cascading input end includes: a clock cascading input end 210; the cascading output end includes: a clock cascading output end 220;

[0126] The clock cascading input end 210 is connected to the input end of the first processing unit 110 and is used to connect the clock cascading output end 220 of another cascaded synchronization device 100;

[0127] The clock cascading output end 220 is connected to the output end of the first processing unit 110.

[0128] In the embodiments of the present disclosure, each synchronization device 100 can be connected to N (N>1, for example, N = 4 or N = 6, etc.) GPUs. That is, at least N input ends of the first processing unit 110 and the AND operation unit 120 in each synchronization device 100 are included. In the embodiments of the present disclosure, each synchronization device 100 can also be used to cascade other synchronization devices 100. Specifically, the input end of the first processing unit 110 can be N + 1, where N input ends are respectively used to connect different GPUs, and the other input end can be used as the above clock cascading input end 210 to cascade with other synchronization devices 100, that is, to connect the output end of the first processing unit 110 of another synchronization device 100. The output end of the first processing unit 110 is then divided into N + 1 branches, where N branches are respectively connected to the reference clock input ends of multiple GPUs, and the other branch is used as the above clock cascading output end 220 to connect to the clock cascading input end 210 of another synchronization device 100. In this way, the synchronization of the reference clocks of multiple synchronization devices 100 can be achieved.

[0129] In some embodiments, when the clock cascading input end 210 is not connected to the clock cascading output end 220 of the other synchronization device 100, the clock cascading input end 210 floats; and / or

[0130] When the clock cascading output end 220 floats and is not connected to the clock cascading input end 210 of the other synchronization device 100, the clock cascading output end 220 floats.

[0131] In addition, in some other embodiments, since the above-mentioned clock cascade input terminal 210 is the input terminal of the first processing unit 110, therefore, in the case of not cascading other synchronization devices, the clock cascade input terminal 210 can be grounded or connected to a low level to provide a logic "0". In this way, the output signal of the first processing unit 110 is not affected by this terminal.

[0132] In some embodiments, as Figure 4 shown, the cascade input terminal includes: a first synchronization input terminal 310; the cascade output terminal includes: a first synchronization output terminal 320;

[0133] The first synchronization input terminal 310 is connected to the input terminal of the second processing unit 120 and is used to connect the first synchronization output terminal 320 of another cascaded synchronization device 100.

[0134] The first synchronization output terminal 320 is connected to the output terminal of the second processing unit 120.

[0135] In some embodiments, the cascade input terminal further includes: a second synchronization input terminal 330, and the cascade output terminal further includes: a second synchronization output terminal 340;

[0136] Both the second synchronization input terminal 330 and the second synchronization output terminal 340 are connected to the synchronization signal input terminals of the multiple GPUs, and the second synchronization input terminal 330 is used to connect the second synchronization output terminal 340 of another cascaded synchronization device 100, or is used to connect the first synchronization output terminal 320 of the synchronization device 100.

[0137] In the embodiments of the present disclosure, the synchronization of the frame synchronization output signal o_fsync of multiple cascaded synchronization devices 100 is achieved through two groups of input and output terminals.

[0138] As described above Figure 4 shown, since in the case of cascading multiple synchronization devices, each synchronization device needs to be connected to the frame signal synchronization of all GPUs. Therefore, in logical operation, the frame synchronization output signals o_fsync of all GPUs need to be ANDed. The connection relationship between the second processing units 120 of two adjacent synchronization devices 100 is: the output terminal of the upper-level second processing unit 120 is connected to the input terminal of the lower-level second processing unit 120, so as to realize the cascading of multiple second processing units 120. In this way, the output signal of the second processing unit 120 of the last synchronization device 100 is the final frame synchronization signal and is used to be provided to the synchronization signal input terminals of all GPUs.

[0139] Therefore, for a synchronization device 100, the synchronization signal input ends of multiple GPUs are connected together, and the second synchronization input end 330 and the second synchronization output end 340 are respectively led out. In addition to connecting to the synchronization signal output ends of multiple GPUs at the same level among the multiple input ends of the arithmetic unit, there is an additional input end as the first synchronization input end 310; the output end of the second processing unit 120 serves as the first synchronization output end 320. When other synchronization devices 100 are cascaded, the first synchronization output end 320 of the upper-level synchronization device 100 is connected to the first synchronization input end 310 of the lower-level synchronization device 100. When there is no cascaded lower-level synchronization device 100, the first synchronization output end 320 is connected to the second synchronization input end 330 of the same level, that is, the output end of the second processing unit 120 is connected to the synchronization signal input ends of each GPU. The second synchronization output end 340 of the synchronization device 100 at the same level is used to be connected to the second synchronization input end 330 of the upper-level synchronization device 100. In this way, the second processing units 120 of multiple cascaded synchronization devices can be cascaded, so that the frame synchronization output signals o_fsync of all GPUs participate in the AND operation, and the final AND operation result is provided to each GPU as the valid synchronization signal i_fsync to indicate frame synchronization.

[0140] For the specific connection method of cascading multiple synchronization devices 100 above, reference can be made to Figure 5 , Figure 5 which shows the cascading situation of 3 synchronization devices 100 and is specifically described as follows:

[0141] In some embodiments, taking the first synchronization device 100(1) in Figure 5 as an example, in the case where no other synchronization device is cascaded on the first side of the synchronization device 100 and another synchronization device is cascaded on the second side:

[0142] Both the first synchronization input end 310 and the second synchronization output end 340 are floating;

[0143] The first synchronization output end 320 is connected to the first synchronization input end 310 of the other synchronization device 100(2);

[0144] The second synchronization input end 330 is connected to the second synchronization output end 340 of the other synchronization device 100(2).

[0145] In some embodiments, taking Figure 5 the second synchronization device 100(2) in Figure 5 as an example, in the case where the first synchronization device (such as the synchronization device 100(1) shown in Figure 5In the case of the synchronization device 100(3) shown:

[0146] The first synchronization input terminal 310 is connected to the first synchronization output terminal 320 of the first synchronization device 100(1);

[0147] The second synchronization output terminal 340 is connected to the second synchronization input terminal 330 of the first synchronization device 100(1);

[0148] The first synchronization output terminal 320 is connected to the first synchronization input terminal 310 of the second synchronization device 100(3);

[0149] The second synchronization input terminal 330 is connected to the second synchronization output terminal 340 of the second synchronization device 100(3).

[0150] Wherein, the first synchronization device 100(1) and the second synchronization device 100(3) and the synchronization device 100(2) may be the same synchronization device 100.

[0151] In some embodiments, taking Figure 5 the synchronization device 100(3) as an example, when another synchronization device (such as Figure 5 the synchronization device 100(2) shown) is cascaded on the first side of the synchronization device 100(3), and there is no other synchronization device cascaded on the second side:

[0152] The first synchronization input terminal 310 is connected to the first synchronization output terminal 320 of the other synchronization device 100(2);

[0153] The second synchronization output terminal 340 is connected to the second synchronization input terminal 330 of the other synchronization device 100(2);

[0154] The first synchronization output terminal 320 is connected to the second synchronization input terminal 330.

[0155] In addition, Figure 5 It is also shown in that the clock cascade input terminals 210 of each synchronization device 100 are sequentially connected to the clock cascade output terminals 220 of the adjacent stage, thereby realizing the cascade of the effective clock signal i_refclk.

[0156] Combining the above three cases, in the case of cascading multiple synchronization devices 100, as the first-stage synchronization device 100(1), there is no other cascaded end at its first synchronization input terminal 310, so it can be floating or set to "1" (providing a high level); the first synchronization output terminal 320 is connected to the first synchronization input terminal 310 of the next-stage synchronization device 100(2); the second synchronization input terminal 330 is connected to the second synchronization output terminal 340 of the next-stage synchronization device 100(2).

[0157] As the intermediate-level synchronization device 100(2), its first synchronization input terminal 310 is connected to the first synchronization output terminal 320 of the upper-level synchronization device 100(1), its first synchronization output terminal 320 is connected to the first synchronization input terminal 310 of the lower-level synchronization device 100(3), its second synchronization input terminal 330 is connected to the second synchronization output terminal 340 of the lower-level synchronization device 100(3), and its second synchronization output terminal 340 is connected to the second synchronization input terminal 330 of the upper-level synchronization device 100(1). It can be understood that if there are more levels of synchronization devices 100 connected to each other, the connection methods of the intermediate-level synchronization devices 100 are the same as those of the above synchronization device 100(2).

[0158] As the last-level synchronization device 100(3), its own first synchronization output terminal 320 is connected to the second synchronization input terminal 330, so as to provide the output signal of the AND operation unit 120 to the synchronization signal input terminals of each GPU, and further transmit it to the synchronization signal input terminals of each GPU of other cascaded synchronization devices 100 through the second synchronization output terminal 340. Thus, the input signals of the synchronization signal input terminals of all GPUs are synchronized to the above effective synchronization signal i_fsync.

[0159] In some embodiments, as Figure 6 shown, when there are no other cascaded synchronization devices at the end of the synchronization device 100:

[0160] Both the first synchronization input terminal 310 and the second synchronization output terminal 340 are floating;

[0161] The first synchronization output terminal 320 is connected to the second synchronization input terminal 330.

[0162] In this way, the output terminal of the second processing unit 120 of the synchronization device 100 is connected to the synchronization signal input terminals of multiple GPUs, so as to perform an AND operation on the frame synchronization output signals o_fsync of multiple GPUs and provide them to each GPU as the effective synchronization signal i_fsync. The first synchronization input terminal 310 and the second synchronization output terminal 340 do not require additional signals and can be floating. In another embodiment, since the first synchronization input terminal 310 is the input terminal of the second processing unit 120, it can also be set to "1", for example, connected to the power supply terminal to provide a high level as an input signal.

[0163] It is understandable that, in the case where other synchronization devices are not cascaded, the input and output of the first processing unit 110 are also limited to the synchronization device itself, so the clock cascade input terminal 210 and the clock cascade output terminal 220 can be floating. In another embodiment, since the clock cascade input terminal 210 is an input terminal of the first processing unit 110, in the case where it is not cascaded, the input terminal can also be set to logic "0" by grounding or connecting to a low-level power supply terminal.

[0164] In some embodiments, as described above Figure 4 As shown, the synchronization device 100 further includes:

[0165] One or more synchronization indicating units 350; the synchronization indicating unit 350 is connected to the synchronization state output terminal of the GPU; the synchronization state output terminal is used for the GPU to indicate that the GPU is in a clock synchronization state based on the received valid clock signal i_refclk.

[0166] When the GPU receives a valid clock signal i_refclk, or receives a valid clock signal i_refclk and processes it through internal logic to determine that the GPU has completed clock synchronization, it can output a clock synchronization state signal o_sync_st through the above-mentioned synchronization state output terminal. This signal can be provided to the above-mentioned synchronization indication unit 350, thereby indicating that the GPU has achieved clock synchronization and can work normally to process frame data of subsequent images.

[0167] Exemplarily, the synchronization indication unit 350 may be an indicator light, so that the user can monitor the synchronization status of the GPU during testing or use. In addition, the synchronization indication unit 350 may also be a status signal output terminal, which is used to output the clock synchronization status signal o_sync_st of each GPU to other devices, such as test equipment, display equipment or other control equipment, so as to notify the other party that the GPU has completed clock synchronization and can perform subsequent frame data processing.

[0168] The present disclosure also provides the following examples:

[0169] Figure 7The logical architecture of splicing image processing using the synchronization device provided by the embodiment of the present disclosure and four graphics cards (GPUs) is shown. Here, a complete image is cut into four equal strip areas (Strip1 to Strip4). For example, the picture is divided into four horizontally extending strip areas along the vertical direction of the image. Each Strip is further divided into four block areas of the same size (Block1-1 to Block4-4). For example, the size of each Block is up to 16K, and the entire picture has a total of 16 Blocks. These 16 Blocks can be provided to 16 display units for multi-screen display, and the 16 display units can be spliced ​​together to display the complete picture. If each GPU supports 4-way display, 4 graphics cards are required to output 16 Blocks. In order to ensure that it can be spliced ​​into a complete picture on the real end, it is necessary to synchronize the outputs of multiple graphics cards without requiring additional configuration of the display end to ensure the synchronous display of multiple display units.

[0170] like Figure 7 As shown, four graphics cards (GPU0 to GPU3) are connected to a synchronization device 100. Each graphics card communicates with the synchronization device 100 through five signals. The five signals include: o_refclk: synchronization reference output clock, the output of the master card in each group of graphics cards is used for clock synchronization of each display port of the board; i_refclk: synchronization reference input clock, each graphics card in the system uses this clock as a reference clock for synchronous output; o_fsync: frame synchronization output signal, used for synchronization of the frame buffer (FrameBuff) of each graphics card in the multi-card system, ensuring that the frame buffer data of the display engines of all graphics cards are ready; i_fsync: valid synchronization signal, used to notify that the frame buffer of all display engines in the multi-card system is ready and can be output for display; o_sync_st: synchronization status output signal, when the reference clock on the graphics card is synchronized, the signal is set high, thereby controlling the indicator light connected to the synchronization device to light up, to indicate that the corresponding graphics card is in a clock synchronization state.

[0171] The logic of the synchronization device 100 is relatively simple, and is composed of input and output signals and signal processing digital logic. The definition of the signal connected to the graphics card synchronization interface is the same. In addition, two sets of input and output signals are provided to realize the cascading of the synchronization device 100, thereby realizing more expansion. Among them, o_ext_i_fsync: the input frame synchronization signal of the external output (that is, the signal output by the second synchronization output terminal involved in the above embodiment) is connected to the i_ext_i_fsync of the previous level (that is, the second synchronization input terminal involved in the above embodiment); i_ext_o_fsync: the input frame synchronization signal of the external input (that is, the input signal of the first synchronization input terminal involved in the above embodiment) is connected to the o_ext_o_fsync of the previous level (that is, the first synchronization output terminal involved in the above embodiment); i_ext_o_refclk: the output synchronization reference of the external input The clock (i.e., the input clock signal of the clock cascade input terminal in the above embodiment) is connected to the o_ext_o_refclk (i.e., the clock cascade output terminal in the above embodiment) of the synchronization device at the previous level; i_ext_i_fsync: the input synchronization frame signal of the external input (i.e., the input signal of the second synchronization input terminal in the above embodiment); o_ext_o_fsync: the frame synchronization signal of the external output (i.e., the input signal of the first synchronization output terminal in the above embodiment); o_ext_o_refclk: the output synchronization reference clock of the external output (i.e., the output signal of the clock cascade output terminal in the above embodiment). There are two signal digital processing logics. Logic01 is an OR operation unit, which realizes that any graphics card is used as the main card, and provides o_refclk output to each secondary card and the i_refclk port of the main card itself. Logic02 is an AND operation unit, which outputs a valid signal only when each input signal is valid (high level, logic "1").

[0172] The configuration logic of the synchronization device 100 provided in the embodiment of the present disclosure for connecting each GPU can be referred to Figure 8 . Specifically, Figure 8 This is the process of configuring multi-card synchronization. In a system, there is only one graphics card as the main card, and the other graphics cards are slave cards. The slave cards receive the reference clock signal sent by the main card to achieve clock synchronization. The specific steps are as follows:

[0173] 1. Configure the secondary card

[0174] Step S701, initializing the display system;

[0175] Step S702, configuring the currently configured graphics card to secondary card mode;

[0176] Step S703, select external i_refclk as the reference clock;

[0177] Step S704: After the clock is stably synchronized, an o_sync_st signal may be output to light up the indicator light of the synchronization device (ie, the synchronization indicating unit involved in the above embodiment);

[0178] Step S705: If there is an image to be output, after the frame buffer is ready, an o_fsync signal is output to inform other graphics cards that the data is ready;

[0179] Step S706, detect the i_fsync signal. If the data of all graphics cards are ready, the signal is valid. The display engine of the current graphics card fetches data from the frame buffer and sends it to the display end, then clears the o_fsync signal to synchronously prepare the data of the next frame.

[0180] Thereafter, the above steps S705 and S706 are looped to realize synchronous display of the video or multiple pictures.

[0181] 2. Configure the primary card

[0182] Step S711, initializing the display system;

[0183] Step S712, configuring the current graphics card to be in master card mode;

[0184] Step S713, outputting a synchronous reference clock to the o_refclk pin;

[0185] Step S714, select the external i_refclk as the reference clock; at this time, the synchronous reference input clock i_refclk signal actually comes from the synchronous reference output clock o_refclk signal output by the graphics card itself;

[0186] Step S715: After the clock is stabilized and synchronized, an o_sync_st signal may be output to light up the indicator light of the synchronization device;

[0187] Step S716: If there is a picture to be output, after the frame buffer is ready, an o_fsync signal is output to inform other secondary cards that the data is ready;

[0188] Step S717, detect the i_fsync signal. If the data of all graphics cards are ready, the signal is valid, the display engine frame buffer of the main card fetches data and sends it to the display end, and then clears the o_fsync signal to synchronously prepare the data of the next frame.

[0189] Then, the above steps S716 and S717 are looped to realize synchronous display of video or multiple pictures.

[0190] It should be noted that in the above configuration process, the secondary card can be configured first, and the above steps S701 to S703 are executed for each secondary card to achieve initialization. Then the primary card is configured, and the above steps 711 to S713 are executed for the primary card to complete initialization.

[0191] After that, each graphics card can execute the subsequent steps simultaneously and enter the step loops of subsequent steps S705 / S706 and steps S716 / S717 synchronously.

[0192] Figure 9 The structural schematic diagram of the synchronization device provided by this embodiment of the present disclosure applied to a display system is shown here. Taking the example that a single synchronization device supports 4 connections (4-channel signals of a single card are automatically synchronized internally) and each channel supports 4 outputs. The image to be displayed is cut into 16 sub-images of the same size and displayed on 16 display units (each display unit can be an independent display screen or an area on a display screen). The display interfaces (DE0 to DE3) of the graphics cards (GPU0 to GPU3) are connected to the display, and the synchronization interface (SYNC) is connected to the synchronization device 100. According to specific implementations, the synchronization device 100 supports directly obtaining power from the motherboard PCIE (Peripheral Component Interconnect Express) or providing a separate interface for power access.

[0193] The above synchronization device provided by this embodiment of the present disclosure is simple to implement, does not occupy too many logic units of the graphics card, does not increase excessive chip production costs, is simple to deploy, and does not require additional dedicated equipment, which can reduce costs. And it can enrich the features of the graphics card and enhance the product competitiveness.

[0194] Based on the same inventive concept, this embodiment of the present disclosure also provides an image processing system, as Figure 10 shown. The image processing system 900 includes:

[0195] One or more of the above-mentioned synchronization devices 100;

[0196] Multiple GPUs, and each of the synchronization devices 100 is connected to one or more of the GPUs (GPU0 to GPU3 as shown in the figure).

[0197] In some embodiments, the image processing system 900 further includes:

[0198] A display module 910, connected to the multiple GPUs;

[0199] Based on the valid clock signal and valid synchronization signal output by the synchronization device 100, the GPU outputs the frame data to be displayed to the display module.

[0200] The display module 910 displays the spliced screen based on the frame data output by the multiple GPUs.

[0201] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities or similarities can be referred to each other. The description of the above embodiments of the image processing system is similar to the description of the corresponding embodiments of the above synchronization device, and has beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the image processing system in the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.

[0202] In the embodiments of the present disclosure, the GPU involved may be an architecture independent of other processors such as the CPU, and may also be referred to as a "graphics card". Figure 11 It is a schematic diagram of the composition structure of a GPU involved in the embodiments of the present disclosure. As Figure 11 shown, the hardware entities of the GUP800 include: a GPU chip 801, a video memory 802, and an interface unit 803;

[0203] Among them, the GPU chip 801 is the core component of the GPU 800, and is used to execute graphics rendering tasks. Exemplarily, it includes vertex processing, rasterization, and pixel shading, etc.;

[0204] The video memory 802 is a dedicated storage medium for the GPU, and can be a memory for storing graphic data, such as storing texture information, vertex information, etc. The types of video memory can include the Graphics Double Data Rate (GDDR) series and the High Bandwidth Memory (HBM), etc.;

[0205] The interface unit 803 is used to connect to the interface of the display or other display devices of the electronic device, such as the High Definition Multimedia Interface (HDMI), the DisplayPort, the Digital Visual Interface (DVI), the Video Graphics Array (VGA), etc.

[0206] In addition, in some embodiments, the above GPU may further include a cooling system, a power interface, and buses and interfaces with other systems in the electronic device (such as the CPU and the storage system, etc.). Moreover, in some embodiments, the GPU may further include specific hardware accelerators, such as Tensor Cores for artificial intelligence and deep learning, or Ray Tracing Cores (RT) for real-time ray tracing, etc.

[0207] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present disclosure, the magnitudes of the serial numbers of the above steps / processes do not mean the sequence of execution is prior or subsequent. The execution sequence of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure above are only for description and do not represent the superiority or inferiority of the embodiments.

[0208] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0209] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0210] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0211] In addition, each functional unit in the embodiments of the present disclosure may be fully integrated into one processing unit, or each unit may be separately regarded as one unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0212] As described above, the above are only the embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure.

Claims

1. A synchronization device for a graphics processing unit GPU, characterized in that: The synchronization device comprises: A first processing unit, configured to receive a preset valid clock signal and output the valid clock signal to at least one GPU; a second processing unit, configured to receive a first frame synchronization signal of the at least one GPU, and output a second frame synchronization signal to the at least one GPU; wherein the second frame synchronization signal is a valid synchronization signal when the first frame synchronization signals all indicate that the GPU is in a frame synchronization state; the valid synchronization signal is used to indicate that the at least one GPU is in a frame synchronization state; the frame synchronization state indicates that the processing of the current frame data has been completed; The GPU outputs the frame data to be displayed based on the valid clock signal in a state where the second frame synchronization signal is the valid synchronization signal.

2. The synchronization device according to claim 1, characterized in that: The first processing unit comprises: An OR operation unit, comprising a plurality of input terminals, and configured to perform an OR operation on signals received by the plurality of input terminals; At least part of the multiple input terminals of the OR operation unit are respectively used to connect to the GPU; Wherein, any of the GPUs inputs the valid clock signal to the OR operation unit, or a preset external input inputs the valid clock signal to the OR operation unit through the input terminal; The output end of the OR operation unit is connected to the at least one GPU, and is used to provide the effective clock signal to the at least one GPU.

3. The synchronization device according to claim 2, characterized in that: The multiple input terminals of the OR operation unit are respectively connected to pull-down resistors; the pull-down resistors are used to set the input signals of the input terminals to logic 0 when the valid clock signal is not received.

4. The synchronization device according to claim 1, characterized in that: The first processing unit comprises: A signal selection unit, comprising a plurality of input terminals, for selecting a valid clock signal from the signals received by the plurality of input terminals as an output signal; At least part of the plurality of input terminals of the signal selection unit are respectively used to connect to the GPU; Wherein, any of the GPUs inputs the valid clock signal to the signal selection unit, or a preset external input inputs the valid clock signal to the signal selection unit through the input terminal; The output end of the signal selection unit is connected to the at least one GPU, and is used to provide the effective clock signal to the at least one GPU.

5. The synchronization device according to claim 1, characterized in that: The second processing unit comprises: AND unit; The AND operation unit includes a plurality of input terminals; at least part of the plurality of input terminals of the AND operation unit are respectively used to connect to the GPU and receive the first frame synchronization signal provided by the GPU; Among them, the first frame synchronization signal of high level is the valid synchronization signal.

6. The synchronization device according to claim 5, characterized in that: The plurality of input terminals of the AND operation unit are respectively connected to pull-up resistors; the pull-up resistors are used to set the input terminals not connected to the GPU to logic 1.

7. The synchronization device according to claim 1, characterized in that: The second processing unit comprises: Counting unit; The counting unit comprises a plurality of input terminals; at least part of the plurality of input terminals of the counting unit are respectively used to connect to the GPU and receive the first frame synchronization signal provided by the GPU; The counting unit is used for counting valid synchronization signals in the received first frame synchronization signal, and outputting the second frame synchronization signal to the at least one GPU when the count value is consistent with the number of the received first frame synchronization signals.

8. The synchronization device according to any one of claims 1 to 7, characterized in that: The synchronization device includes N GPU interfaces; the GPU interfaces include at least N first connection ends connected to the first processing unit and N second connection ends connected to the second processing unit; wherein N is an integer greater than or equal to 2; The synchronization device is connected to the M GPUs through the GPU interface, and is used to provide the effective clock signal to the M GPUs through the first processing unit, and provide the second frame synchronization signal to the M GPUs through the second processing unit; Wherein, M is a positive integer less than or equal to N.

9. The synchronization device according to any one of claims 1 to 7, characterized in that: The synchronization device also includes: A cascade input terminal and a cascade output terminal; The cascade input terminal is connected to the cascade output terminal of another synchronization device in the cascade; The cascade output terminal and the cascade input terminal are used to transmit the valid clock signal and / or the valid synchronization signal between the plurality of cascaded synchronization devices.

10. The synchronization device according to claim 9, characterized in that: The cascade input terminal includes: a clock cascade input terminal; the cascade output terminal includes: a clock cascade output terminal; The clock cascade input terminal is connected to the input terminal of the first processing unit and is used to connect to the clock cascade output terminal of another synchronization device in the cascade; The clock cascade output terminal is connected to the output terminal of the first processing unit.

11. The synchronization device according to claim 10, characterized in that: When the clock cascade input terminal is not connected to the clock cascade output terminal of the other synchronization device, the clock cascade input terminal is floating; and / or When the clock cascade output terminal is floating and is not connected to the clock cascade input terminal of another synchronization device, the clock cascade output terminal is floating.

12. The synchronization device according to claim 9, characterized in that: The cascade input terminal includes: a first synchronization input terminal; the cascade output terminal includes: a first synchronization output terminal; The first synchronization input terminal is connected to the input terminal of the second processing unit and is used to connect to the first synchronization output terminal of another cascaded synchronization device; The first synchronous output terminal is connected to the output terminal of the second processing unit.

13. The synchronization device according to claim 12, characterized in that: The cascade input terminal further includes: a second synchronization input terminal, and the cascade output terminal further includes: a second synchronization output terminal; The second synchronization input terminal and the second synchronization output terminal are both connected to the at least one GPU for transmitting the valid synchronization signal, and the second synchronization input terminal is used to connect to the second synchronization output terminal of another cascaded synchronization device, or to connect to the first synchronization output terminal of the synchronization device.

14. The synchronization device according to claim 13, characterized in that: In the case where no other synchronization device is cascaded to the first side of the synchronization device, and another synchronization device is cascaded to the second side: The first synchronization input terminal and the second synchronization output terminal are both floating; The first synchronization output terminal is connected to the first synchronization input terminal of the other synchronization device; The second synchronization input terminal is connected to the second synchronization output terminal of the other synchronization device.

15. The synchronization device according to claim 13, characterized in that: In the case where a first synchronization device is cascaded with a first synchronization device on the first side and a second synchronization device is cascaded with a second synchronization device on the second side: The first synchronization input terminal is connected to the first synchronization output terminal of the first synchronization device; The second synchronization output terminal is connected to the second synchronization input terminal of the first synchronization device; The first synchronization output terminal is connected to the first synchronization input terminal of the second synchronization device; The second synchronization input terminal is connected to the second synchronization output terminal of the second synchronization device; The first synchronization device and the second synchronization device are the same as the synchronization device.

16. The synchronization device according to claim 13, characterized in that: In the case where another synchronization device is cascaded to the first side of the synchronization device, and no other synchronization device is cascaded to the second side: The first synchronization input terminal is connected to the first synchronization output terminal of the other synchronization device; The second synchronization output terminal is connected to the second synchronization input terminal of the other synchronization device; The first synchronization output terminal is connected to the second synchronization input terminal.

17. The synchronization device according to claim 13, characterized in that: In the case where the synchronization device is not cascaded with other synchronization devices: The first synchronization input terminal and the second synchronization output terminal are both floating; The first synchronization output terminal is connected to the second synchronization input terminal.

18. The synchronization device according to any one of claims 1 to 17, characterized in that: The synchronization device also includes: One or more synchronization indicating units; the synchronization indicating unit is used to receive a clock synchronization state signal sent by the GPU based on the valid clock signal, indicating that the GPU is in a clock synchronization state.

19. An image processing system, characterized in that: include: One or more cascaded synchronization devices as claimed in any one of claims 1 to 18; There are multiple GPUs, and each of the synchronization devices is connected to one or more of the GPUs.

20. The image processing system according to claim 19, characterized in that: Also includes: A display module connected to the multiple GPUs; The GPU outputs frame data to be displayed to the display module based on the valid clock signal and the valid synchronization signal output by the synchronization device; The display module displays a spliced ​​picture based on the frame data output by the multiple GPUs.

Citation Information

Patent Citations

  • Synchronous display device, stacking splicing display system and synchronous display method thereof

    CN101763840A