Synchronous communication device, method and equipment based on multiple coprocessors and storage medium
Through a synchronous communication device based on a multi-coprocessor, the distribution of the first-level buffer and multiple second-level buffers and the parallel operation of the coprocessors is solved, and the communication efficiency and CPU utilization are improved.
Patent Information
- Application Number
- CN202510284546.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
When multiple terminal nodes conduct synchronous communication, in the prior art, a single synchronization instruction sends a long period and a long waiting time for data writing and reading of the terminal node, resulting in low communication efficiency.
Using a synchronous communication device based on a multi-coprocessor, the communication instructions are sent to the first-level buffer through the main processor. The cache controller distributes the instructions to multiple second-level buffers. The coprocessor reads and configures them to the corresponding terminal nodes respectively. The synchronous trigger module controls the node communication, and stores the results in the buffer, and finally parses the results by the main processor.
Parallel communication between multiple terminal nodes is realized, which shortens the instruction sending and reading time, improves communication efficiency, and reduces the real-time requirements for the CPU.
Smart Images

Figure CN120256374A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a synchronous communication device, method, equipment, and storage medium based on multiple coprocessors. Background Art
[0002] The main processor sequentially writes the data and instructions to be sent into each terminal node, and uses the backplane synchronous trigger function to realize the synchronous communication of data at multiple sites. After the communication is completed, the data results are read out sequentially. In this way, since the data and instructions to be sent of each terminal node are serially written, the single synchronous instruction sending cycle is long, and most of the time is spent waiting for the writing and reading of data of other terminal nodes.
[0003] In view of this, how to shorten the instruction sending cycle during synchronous communication among multiple terminal nodes has become an urgent technical problem to be solved. Summary of the Invention
[0004] In view of this, an object of the present disclosure is to provide a synchronous communication device, method, equipment, and storage medium based on multiple coprocessors to solve or partially solve the above technical problems.
[0005] Based on the above object, a first aspect of the present disclosure provides a synchronous communication device based on multiple coprocessors, the device includes: a main processor, a first-level buffer, a cache controller, a synchronous trigger module, a first coprocessor, a first second-level buffer, a first terminal node, a second coprocessor, a second second-level buffer, and a second terminal node;
[0006] The main processor is configured to send a communication instruction to the first-level buffer;
[0007] The cache controller is configured to distribute the communication instruction from the first-level buffer to the first second-level buffer and the second second-level buffer;
[0008] The first coprocessor is configured to read the communication instruction from the first second-level buffer and configure the communication instruction to the first terminal node;
[0009] The second coprocessor is configured to read the communication instruction from the second second-level buffer and configure the communication instruction to the second terminal node;
[0010] The synchronous trigger module is configured to control the first terminal node and the second terminal node to perform synchronous communication according to the communication instruction to obtain a communication result;
[0011] The first coprocessor is configured to read the communication result from the first terminal node and store the communication result in the first secondary buffer;
[0012] The second coprocessor is configured to read the communication result from the second terminal node and store the communication result in the second secondary buffer;
[0013] The cache manager is configured to collect the communication results in the first secondary buffer and the second secondary buffer into the primary buffer;
[0014] The main processor is configured to collect the communication results in the primary buffer into the host computer, and use the host computer to parse and process the communication results to obtain a parsing result.
[0015] Based on the same inventive concept, a second aspect of the present disclosure proposes a synchronous communication method, which is applied to the synchronous communication device based on multiple coprocessors described in the first aspect; the device includes: a main processor, a primary buffer, a cache controller, a synchronization trigger module, a first coprocessor, a first secondary buffer, a first terminal node, a second coprocessor, a second secondary buffer, and a second terminal node; the method includes:
[0016] The main processor sends a communication instruction to the primary buffer;
[0017] The cache controller distributes the communication instruction from the primary buffer to the first secondary buffer and the second secondary buffer;
[0018] The first coprocessor reads the communication instruction from the first secondary buffer and configures the communication instruction into the first terminal node;
[0019] The second coprocessor reads the communication instruction from the second secondary buffer and configures the communication instruction into the second terminal node;
[0020] The synchronization trigger module controls the first terminal node and the second terminal node to perform synchronous communication according to the communication instruction to obtain a communication result;
[0021] The first coprocessor reads the communication result from the first terminal node and stores the communication result in the first secondary buffer;
[0022] The second coprocessor reads the communication result from the second terminal node and stores the communication result in the second secondary buffer;
[0023] The cache manager collects the communication results in the first secondary cache and the second secondary cache into the primary cache;
[0024] The main processor collects the communication results in the primary cache into the host computer, and uses the host computer to perform parsing processing on the communication results to obtain a parsing result.
[0025] Based on the same inventive concept, a third aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor. When the processor executes the computer program, the method described above is implemented.
[0026] Based on the same inventive concept, a fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method described above.
[0027] As can be seen from the above, the present disclosure provides a synchronous communication device, method, device, and storage medium based on multiple coprocessors. The main processor sends a communication instruction to the primary cache. The cache controller distributes the communication instruction from the primary cache to the first secondary cache and the second secondary cache. The first coprocessor reads the communication instruction from the first secondary cache and configures the communication instruction into the first terminal node. The second coprocessor reads the communication instruction from the second secondary cache and configures the communication instruction into the second terminal node. The synchronization trigger module controls the first terminal node and the second terminal node to perform synchronous communication according to the communication instruction to obtain a communication result. The first coprocessor reads the communication result from the first terminal node and stores the communication result in the first secondary cache. The second coprocessor reads the communication result from the second terminal node and stores the communication result in the second secondary cache. The cache manager collects the communication results in the first secondary cache and the second secondary cache into the primary cache. The main processor collects the communication results in the primary cache into the host computer, and uses the host computer to perform parsing processing on the communication results to obtain a parsing result. In this way, the first coprocessor and the second coprocessor can simultaneously configure the corresponding terminal nodes, thereby shortening the instruction sending time. The first coprocessor and the second coprocessor can simultaneously read the communication results from the corresponding terminal nodes, thereby shortening the instruction reading time, and there is no need to wait for the writing and reading of other terminal nodes. Description of the Drawings
[0028] To more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0029] Figure 1 Schematic structural diagram of a GLink communication device in related technologies;
[0030] Figure 2 Flowchart of a GLink communication method in related technologies;
[0031] Figure 3 Schematic structural diagram of a synchronous communication device based on multiple coprocessors according to an embodiment of the present disclosure;
[0032] Figure 4 Schematic structural diagram of a data communication device according to an embodiment of the present disclosure;
[0033] Figure 5 Flowchart of a data communication method according to an embodiment of the present disclosure;
[0034] Figure 6 Schematic diagram of a synchronous communication device based on multiple coprocessors according to another embodiment of the present disclosure;
[0035] Figure 7 Schematic structural diagram of a coprocessor according to an embodiment of the present disclosure;
[0036] Figure 8 Flowchart for an instruction processing module to judge a request according to an embodiment of the present disclosure;
[0037] Figure 9 Flowchart for exchanging data according to an embodiment of the present disclosure;
[0038] Figure 10 Schematic structural diagram of a transceiver process handling module according to an embodiment of the present disclosure;
[0039] Figure 11 Flowchart for a transceiver process handling module to control messages according to an embodiment of the present disclosure;
[0040] Figure 12 Timing diagram of data transceiver according to an embodiment of the present disclosure;
[0041] Figure 13 Schematic structural diagram of a cache controller according to an embodiment of the present disclosure;
[0042] Figure 14 Schematic structural diagram of a message control module according to an embodiment of the present disclosure;
[0043] Figure 15 Flow chart of data processing by the message control module according to an embodiment of the present disclosure;
[0044] Figure 16 Flow chart of data reading and writing by the message reading and writing module according to an embodiment of the present disclosure;
[0045] Figure 17 Structural schematic diagram of the synchronization trigger module according to an embodiment of the present disclosure;
[0046] Figure 18 Flow chart of synchronous communication by the clock synchronization module according to an embodiment of the present disclosure;
[0047] Figure 19 Flow chart of dynamically adjusting the local pulse period according to an embodiment of the present disclosure;
[0048] Figure 20 Schematic diagram of the trigger signal according to an embodiment of the present disclosure;
[0049] Figure 21 Flow chart of the host computer forming and sending a message queue according to an embodiment of the present disclosure;
[0050] Figure 22 Schematic diagram of the timestamp for the coprocessor to control GLink communication according to an embodiment of the present disclosure;
[0051] Figure 23 Structural schematic diagram of the board according to an embodiment of the present disclosure;
[0052] Figure 24 Flow chart of the synchronous communication method based on multiple coprocessors according to an embodiment of the present disclosure;
[0053] Figure 25 Structural schematic diagram of the electronic device according to an embodiment of the present disclosure.
[0054] Description of reference numerals:
[0055] 100. Synchronous communication device based on multiple coprocessors; 110. Main processor; 120. Primary cache; 130. Cache controller; 140. Synchronization trigger module; 151. First coprocessor, 1511. Register module, 1512. Transceiver process handling module, 1513. Instruction processing module, 152. Second coprocessor; 161. First secondary cache, 162. Second secondary cache; 171. First terminal node, 172. Second terminal node. Detailed implementation manners
[0056] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0057] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0058] Based on the description of the background technology, the main processor (Central Processing Unit, abbreviated as CPU) directly writes the data and instructions to be sent into the memory of each GLink terminal node in sequence through PCIe, and then uses the backplane synchronous trigger function to achieve synchronous communication of data at multiple sites. After the communication is completed, the data results are read out through PCIe in sequence.
[0059] Figure 1 It is a schematic structural diagram of a GLink communication device in the related art. As Figure 1 shown, the main processor CPU controls GLink communication IP1, GLink communication IP2, GLink communication IP3 and GLink communication IP4 respectively.
[0060] Figure 2 It is a flowchart of a GLink communication method in the related art. As Figure 2 shown, when the GLink communication includes GLink1, GLink2 and GLink3, the specific steps of the GLink communication are as follows: the main processor configures the GLink1 communication data; the main processor configures the GLink2 communication data; the main processor configures the GLink3 communication data; start all GLink transmissions; the main processor reads the GLink1 communication result; the main processor reads the GLink2 communication result; the main processor reads the GLink3 communication result.
[0061] It has the following disadvantages:
[0062] 1. The read / write rate of the GLink terminal node interface is low, about 10 Mbyte / s, resulting in low utilization rate of the PCIe bandwidth.
[0063] 2. The data and instructions to be sent by each GLink terminal node are serially written, resulting in a long single synchronous instruction sending cycle, and most of the time is spent waiting for the writing and reading of data by other terminal nodes.
[0064] 3. It is necessary to actively start a synchronous trigger regularly after the data and instructions of each terminal node are written.
[0065] 4. The host computer CPU needs to frequently participate in GLink communication, resulting in the need for a processor and operating system with strong real-time performance.
[0066] However, in some application scenarios (such as synchronous control of multiple servos, etc.), it is necessary to increase the sending and receiving frequency of data instructions.
[0067] As described above, how to shorten the instruction sending cycle during synchronous communication among multiple terminal nodes has become an important research issue.
[0068] Based on the above description, as Figure 3 shown, the synchronous communication device 100 based on multiple coprocessors proposed in this embodiment, the device 100 includes: a main processor 110, a first-level buffer 120, a cache controller 130, a synchronous trigger module 140, a first coprocessor 151, a first second-level buffer 161, a first terminal node 171, a second coprocessor 152, a second second-level buffer 162, and a second terminal node 172;
[0069] The main processor 110 is configured to send communication instructions to the first-level buffer 120;
[0070] The cache controller 130 is configured to distribute the communication instructions from the first-level buffer 120 to the first second-level buffer 161 and the second second-level buffer 162;
[0071] The first coprocessor 151 is configured to read the communication instructions from the first second-level buffer 161 and configure the communication instructions into the first terminal node 171;
[0072] The second coprocessor 152 is configured to read the communication instructions from the second second-level buffer 162 and configure the communication instructions into the second terminal node 172;
[0073] The synchronous trigger module 140 is configured to control the first terminal node 171 and the second terminal node 172 to perform synchronous communication according to the communication instructions to obtain a communication result;
[0074] The first coprocessor 151 is configured to read the communication result from the first terminal node 171 and store the communication result in the first secondary buffer 161;
[0075] The second coprocessor 152 is configured to read the communication result from the second terminal node 172 and store the communication result in the second secondary buffer 162;
[0076] The cache manager 130 is configured to collect the communication results in the first secondary buffer 161 and the second secondary buffer 162 into the primary buffer 120;
[0077] The main processor 110 is configured to collect the communication results in the primary buffer 120 into the host computer, and use the host computer to parse and process the communication results to obtain a parsing result.
[0078] In specific implementation, the present disclosure relates to the field of a new generation of aerospace on-board command-responsive fiber optic control buses, and specifically relates to a fiber optic control bus data acceleration method based on a 32-bit instruction code coprocessor and a secondary cache management mechanism. The system includes multiple coprocessors, a primary memory, a secondary memory, and a cache manager. The primary memory is used to receive control instructions, data sent by the host computer in a specified instruction set format, and store the status results of fiber optic control bus communication; the cache manager moves the control instructions and data to be sent stored in the primary memory to the secondary memory and moves the communication results stored in the secondary memory to the primary memory; the coprocessor takes out the control instructions and data from the secondary memory, configures the fiber optic communication bus controller according to the control instructions, and sends data through the fiber optic communication bus controller, and writes the communication results into the secondary memory. Multiple coprocessors are used to parallelly process the time-consuming operations of reading and writing configurations of the fiber optic bus network communication interface. The main processor reduces the CPU occupancy rate and the real-time requirement for the CPU by configuring the communication instruction set and configuring the coprocessor and the cache manager. Compared with the direct control of the fiber optic bus network communication by the main processor CPU, the efficiency of the fiber optic network communication is increased by more than 10 times.
[0079] The present disclosure relates to a new generation of aerospace on-board bus technology (the Glink bus is a command-responsive serial fiber optic bus based on the FC-AE-1553 protocol. The bus meets the requirements of the new generation of on-board communication for high speed, strong real-time performance, integration, etc., and also has extremely high application flexibility). In particular, it is a module that provides the functions of multiple GLink terminal nodes, and can provide multiple independent control flow controllers CtrlNC and control flow terminals CtrlNT.
[0080] After the GLINK bus was launched, during the development of bus applications, in some scenarios, it is necessary to conduct synchronous communication tests on multiple terminal nodes, send control instructions to each terminal node synchronously, and receive the status information of each terminal. However, when the traditional method configures the data to be sent by multiple GLink terminal nodes through a single main processor, it needs to be configured one by one, which cannot meet the requirement of increasing communication frequency. There is an urgent need to develop a method that can configure multiple GLink terminal nodes simultaneously.
[0081] Figure 4 It is a schematic structural diagram of the data communication device according to an embodiment of the present disclosure. As Figure 4 shown, the host CPU is used to send communication instructions to the primary buffer. The buffer controller is used to distribute the communication instructions from the primary buffer to the secondary buffer 1, secondary buffer 2, secondary buffer 3, and secondary buffer 4. The coprocessor 1 configures the communication instructions in the secondary buffer 1 into the GLink communication IP1. The processes of the coprocessor 2, coprocessor 3, and coprocessor 4 configuring the corresponding GLink communication IPs are similar and will not be elaborated here. The hardware synchronization trigger controls the coprocessor 1, coprocessor 2, coprocessor 3, and coprocessor 4 to control the synchronous communication of the GLink communication IP1, GLink communication IP2, GLink communication IP3, and GLink communication IP4 to obtain the communication results. The coprocessor 1 stores the communication results in the GLink communication IP1 into the secondary buffer 1. The processes of the coprocessor 2, coprocessor 3, and coprocessor 4 storing the communication results in the corresponding GLink communication IPs into the corresponding secondary buffers are similar and will not be elaborated here. The buffer controller is also used to store the communication results in the secondary buffer 1, secondary buffer 2, secondary buffer 3, and secondary buffer 4 into the primary buffer.
[0082] Figure 5 It is a flowchart of the data communication method according to an embodiment of the present disclosure. As Figure 5As shown, when the GLink communication includes GLink1, GLink2, and GLink3, the data communication method includes: the main processor packs the communication instruction data to be sent and sends it to the first-level cache through Direct Memory Access (DMA); the cache manager distributes the communication instruction data in the first-level cache to each second-level cache; the coprocessor 1 reads the communication instruction data from the second-level cache 1 and configures the GLink1 communication data; the coprocessor 2 reads the communication instruction data from the second-level cache 2 and configures the GLink2 communication data; the coprocessor 3 reads the communication instruction data from the second-level cache 3 and configures the GLink3 communication data; start all GLink transmissions; the coprocessor 1 reads the GLink1 communication result and stores it in the second-level cache 1; the coprocessor 2 reads the GLink2 communication result and stores it in the second-level cache 2; the coprocessor 3 reads the GLink3 communication result and stores it in the second-level cache 3; the cache manager receives the communication results in each second-level cache into the first-level cache; the main processor receives the communication results in the first-level cache into the host computer through DMA and parses them.
[0083] The present disclosure discloses a GLink node control method, which realizes synchronous and efficient communication of multiple GLink sites through multiple coprocessors and second-level caches. The specific implementation scheme is as follows:
[0084] 1. Use the 32-bit instruction encoding of the coprocessor to process the queue composed of GLink communication message data.
[0085] 2. Can be instantiated as multiple parallel coprocessors according to the application.
[0086] 3. The host computer encodes the multiple Glink instructions to be sent into bytecodes in the manner of coprocessor 32-bit instruction codes and operands.
[0087] 4. The host computer sends the GLink message queue to the first-level cache of the board's Double-Data-Rate Three Synchronous Dynamic Random Access Memory (DDR3) through the DMA of PCIe; and updates the position of the sent message in the first-level cache.
[0088] 5. The cache management module takes out the messages in the DDR3 first-level cache one by one, transports them to the corresponding Bram second-level cache according to the target site of the message, and updates the position of the sent message in the second-level cache.
[0089] 6. The coprocessor fetches the bytecode from the secondary cache, and the prefetch instruction prefetches according to an instruction group of the bytecode. It executes and writes to the memory of Glink, and performs the message exchange operation of the bus according to the clock beat of the synchronization pulse.
[0090] 7. After a group of Glink message groups are sent, the coprocessor writes the next group of Glink messages to be sent into Glink. After writing, it retrieves the result of the previous group of Glink message groups, writes it to the receiving part of the corresponding Bram secondary cache, and updates the write pointer of the receiving part of the Bram secondary cache.
[0091] 8. The cache management module transfers the Glink message sending result to the receiving part of the DDR primary cache according to the write pointer of the receiving part of each Bram secondary cache; and updates the write pointer of the receiving part of the primary cache.
[0092] The host computer reads the Glink message result into the host computer through DMA, and parses each Glink message through the deframing algorithm.
[0093] Through the above embodiments, the first coprocessor and the second coprocessor can simultaneously configure the corresponding terminal nodes, thereby shortening the instruction sending time. The first coprocessor and the second coprocessor can simultaneously read the communication results from the corresponding terminal nodes, thereby shortening the instruction reading time, without waiting for the writing and reading of other terminal nodes.
[0094] In some embodiments, as Figure 6 shown, the first coprocessor 151 includes: a register module 1511, a transceiver process handling module 1512, and an instruction processing module 1513;
[0095] The instruction processing module 1513 is configured to read the communication instruction from the first secondary cache, and parse and process the communication instruction to obtain the instruction configuration position;
[0096] The register module 1511 is configured to configure the communication instruction to the instruction configuration position in the first terminal node 171, and record the read address of the secondary cache in the downward direction and the write address of the secondary cache in the upward direction of the first coprocessor 151;
[0097] The instruction processing module 1513 is configured to determine that the number of messages in the exchange group of the communication instruction is equal to the message sequence number, and send a configuration completion notification to the transceiver process handling module 1512;
[0098] The transceiver process handling module 1512 is configured to perform initialization settings on the first terminal node 171, so that the first terminal node 171 performs synchronous communication under the control of the synchronization trigger module 140;
[0099] The instruction processing module 1513 is configured to read the communication result from the first terminal node 171 in a parallel manner and store the communication result in the first secondary buffer 161 in the upstream direction.
[0100] Specifically, when implemented, Figure 7 is a schematic structural diagram of a coprocessor according to an embodiment of the present disclosure. As Figure 7 shown, the coprocessor includes: a register module 1511, a transceiver process handling module 1512, and an instruction processing module 1513.
[0101] The host computer sends an instruction group to the first-level buffer in the downstream direction (from the host to the board, Host to Card, abbreviated as H2C direction); the cache controller moves the instruction group in the H2C direction first-level buffer to the corresponding H2C direction second-level buffer; the instruction processing module 1513 of the coprocessor prefetches instructions from the second-level buffer, parses the instructions and performs pre-execution (configuring the registers and memories of GLink, setting the data and target sites of GLink communication, etc.), and executes process control to send GLink instructions; at the same time, it parallelly obtains the data results of GLink communication and sends the data results to the corresponding second-level buffer in the upstream direction (from the board to the host, Card to Host, abbreviated as C2H direction).
[0102] When the instruction processing module 1513 prefetches an instruction where the number of messages in the exchange group is equal to the sequence number of the messages in the exchange group, this instruction is the last instruction in the exchange group. At this time, it notifies the transceiver process handling module 1512, and the transceiver process handling module 1512 of the coprocessor starts the initialization settings for the GLink exchange group transmission and enables the input of the synchronization trigger signal; waits for the 1ms timing trigger of the hardware synchronization clock sending module to complete the GLink exchange group transmission process.
[0103] In addition, multiple coprocessors can be instantiated to perform bus data exchange of multiple Glink controllers in a hardware manner in parallel.
[0104] The register module 1511 mainly configures the GLink stack area base address, GLink data area base address, GLink stack size, GLink data area size, and exchange end check mode. At the same time, it records various states of the coprocessor: including the read address of the current H2C secondary buffer, the write address of the C2H secondary buffer, etc.
[0105] Through the above solution, when the instruction processing module 1513 determines that the number of messages in the exchange group of the communication instruction is equal to the message sequence number, it can determine that the communication instruction is the last instruction in the exchange group and send a configuration completion notice to the transceiver process processing module 1512. In this way, it is possible to accurately identify the completion of the communication instruction configuration and notify the transceiver process processing module 1512 after the communication instruction configuration is completed.
[0106] In some embodiments, the instruction processing module 1513 is further configured to:
[0107] Receive a request to be processed and perform judgment processing on the request to be processed;
[0108] In response to determining that the request to be processed is an instruction sending request, read the communication instruction from the first-level and second-level buffer 161 in the downlink direction, write the configuration instruction timestamp into the first-level and second-level buffer 161 in the uplink direction, and configure the communication instruction to the data stack of the first terminal node 171;
[0109] In response to determining that the instruction to be processed is a result reading instruction, read the message instruction group from the first-level and second-level buffer 161 in the uplink direction, write the timestamp and the current state into the first-level and second-level buffer 161 in the uplink direction, and read the communication result from the stack of the first terminal node 171;
[0110] In response to determining that the instruction to be processed is an exchange end instruction, check the stack exchange count of the first terminal node 171;
[0111] In response to determining that the instruction to be processed is an initialization instruction, configure the stack pointer and stack count of the first terminal node 171.
[0112] In specific implementation, the specific content of the instruction processing module 1513 for co-processor instruction prefetching, parsing, and execution mainly includes setting transmission data, reading communication results, initializing the exchange group transmission settings, and querying whether the exchange group exchange is completed.
[0113] Figure 8 This is a flowchart of the instruction processing module for judging requests in the embodiments of the present disclosure. As Figure 8As shown, determine the request type. When the request is a set send data request, the instruction prefetch reads a message instruction group from the secondary cache in the H2C direction, writes the configuration instruction timestamp to the secondary cache in the C2H direction, configures the data stack of GLink, compiles the communication control word according to the configuration instruction, judges the data transfer direction in the configuration instruction, and writes the control word and data or writes the control word. When the request is a read communication result request, the instruction prefetch reads a message instruction group from the secondary cache in the C2H direction, writes the timestamp and the current status to the secondary cache in the C2H direction, reads the GLink stack, reads the GLink control word and data, and writes the stack, control word, data, etc. to the C2H secondary cache. When the request is to check whether the data exchange is completed, check the stack exchange count. When the request is a send initialization request, configure the GLink stack pointer and stack count.
[0114] Set send data: The coprocessor prefetches a communication instruction group from the secondary cache in the H2C direction; writes the configuration instruction in the instruction group and status information such as the current timestamp to the corresponding receiving position of this instruction in the secondary cache in the C2H direction; at the same time, parses the instruction group, parses the sequence number of this message in the GLink communication exchange group and the number of messages in the exchange group, and parses information such as the data length, target site ID number, target sub-address, and data transfer direction of this message; configures this GLink communication message according to the content parsed by the instruction; writes the data instruction of the communication instruction group to the GLink send memory. When parsing the configuration instruction, if it is checked that the sequence number in the exchange group is the same as the number of internal communications in the exchange group, set the exchange group end flag to 1; if the sequence number of the exchange group is greater than the maximum number of sendable exchange groups that can be set, set the exchange group end flag to 1.
[0115] Read communication result: The coprocessor retrieves the corresponding configuration instruction of the communication result to be received from the secondary cache in the C2H direction; parses the configuration instruction to obtain information such as the data length to be read in the communication result; writes the current timestamp and the current status information to the secondary cache in the C2H direction; reads the data required in the communication result, including the description stack, communication control word, and communication data; writes the read communication result to the secondary cache in the C2H direction.
[0116] Check whether the communication data exchange is completed: Query the stack exchange count of GLink communication and judge whether the communication data exchange is completed; if the exchange is completed, set the communication exchange end flag to 1.
[0117] Send initialization: Configure the stack pointer and stack exchange count of the GLink communication exchange group, and after the configuration is completed, clear the communication exchange end flag to 0 and clear the exchange group end flag to 0.
[0118] Figure 9 This is the flowchart of the exchanged data for the embodiments of the present disclosure. AsFigure 9 As shown, the secondary cache in the H2C direction configures GLink according to the instruction parsing content, exchanges control words and exchange data; the secondary cache in the H2C direction sets the communication data corresponding GLink write address pointer in the secondary cache in the C2H direction, message configuration instructions and message identification timestamp; the secondary cache in the C2H direction reads the communication result from GLink corresponding to the GLink read write address pointer in the secondary cache, describes the stack, exchange control words and exchange data.
[0119] Through the above solution, by judging and processing the request to be processed, the request type of the request to be processed can be accurately identified, and then the corresponding request can be executed according to the request type of the request to be processed. When the request to be processed is an instruction sending request, the communication instruction is accurately configured into the first terminal node 171. When the instruction to be processed is a result reading instruction, the communication result in the first terminal node 171 is read into the first secondary cache 161. When the instruction to be processed is an exchange end instruction, the stack exchange number of the first terminal node 171 is checked to accurately identify whether the exchange is over. When the instruction to be processed is an initialization instruction, the stack pointer and stack count of the first terminal node 171 are configured to realize the initialization setting of the first terminal node 171.
[0120] In some embodiments, the transceiver process handling module 1512 is further configured to:
[0121] In response to determining that the readable messages in the first terminal node 171 are greater than a preset message number threshold, receive the communication result of the first terminal node 171;
[0122] In response to determining that there is a communication instruction in the first secondary cache 161 in the downlink direction and the write request flag bit of the first terminal node 171 is in the presence request state, send an instruction issuing request to the instruction processing module 1513, so that the instruction processing module 1513 configures the communication instruction to the first terminal node 171 according to the instruction issuing request;
[0123] In response to determining that the initialization request flag bit is in the presence request state, the transceiver process handling module 1512 performs an initialization setting on the first terminal node 171;
[0124] In response to determining that the exchange end request flag bit is in the presence request state and the timing has ended, check whether the first terminal node 171 has sent the communication result to the first secondary cache 161.
[0125] Specifically, the transceiver process control module 1512 controls the overall operation of the coprocessing and the process of GLink communication, and is used to calculate various address pointers for instruction execution, judge instruction prefetch conditions, result readback conditions, etc.
[0126] Figure 10 This is a schematic structural diagram of the transceiver process handling module according to an embodiment of the present disclosure. As Figure 10 shown, the message flow control (program counter) of the transceiver process handling module includes: GLink exchange group exchange count calculation, GLink exchange group stack pointer calculation, GLink write stack pointer address calculation, GLink write data pointer address calculation, GLink read stack pointer address calculation, GLink read data pointer address calculation, H2C secondary cache read address calculation, C2H secondary cache write address calculation when reading GLink, and C2H secondary cache write address calculation when writing GLink. The message flow control (program counter) of the transceiver process handling module further includes: synchronous trigger input enable control, GLink exchange end judgment, GLink check exchange end requirement judgment, check GLink exchange end timer, GLink set send data requirement judgment, GLink readable message count calculation in GLink, and GLink send initialization request requirement judgment. The transceiver process handling module receives an exchange end interrupt or queries that the exchange count is 0xffff; the check request is set to 1 after the exchange group starts sending and set to 0 after the exchange ends; the timer is started after the check is completed; the write request is set to 0 after the exchange group is set and set to 1 after the exchange ends; it is decremented by one for each read and incremented by the number of message groups after the exchange ends; the send initialization request is set to 1 after the exchange group is set and set to 0 after the message is sent.
[0127] In the transceiver process handling module 1512 of the coprocessor, each address required for calculating data processing is calculated. Each communication message is a fixed-length message. If the communication data volume is insufficient, the empty space is filled with 0. When calculating the address, a fixed length is added each time. When the set space size in the register is reached, the address wraps around to the initial address.
[0128] Figure 11 This is a flowchart of the transceiver process handling module of the present disclosure for controlling messages. As Figure 11 shown, in the coprocessor in the transceiver process handling module, first, the number of readable messages in the GLink memory is judged. When the number of readable messages is almost full, an immediate GLink receive communication result request is initiated, and the communication result is read and executed in the instruction processing module. Then, it is judged whether the write request of the GLink communication, that is, the set send data request, is set to 1. If it is set to 1 and there is data to be sent in the H2C secondary cache, a GLink set send data request is initiated. Then, it is judged whether there is a GLink send initialization request. If there is, a GLink send initialization is initiated. Then, it is judged whether there is a GLink check whether the exchange is over request. If there is and the timer has counted down to the end, it is checked whether the GLink exchange is over. Finally, it is judged whether there is a readable message in the GLink. If there is, a GLink receive communication result request is sent.
[0129] After receiving a communication result, the readable message count in GLink is decremented by 1, the co-processor read message count is updated, the GLink read stack pointer address is updated, the GLink read data pointer address is updated, and the write pointer address when reading GLink in the C2H direction secondary cache is updated. After each piece of transmission data is set in GLink, the co-processor write message count is updated, the set message count within the exchange group is updated, the GLink write stack pointer address is updated, the GLink write data pointer address is updated, the read data pointer address of the H2C direction secondary cache, and the corresponding write pointer address when writing to GLink in the C2H direction secondary cache is updated. After GLink completes the transmission initialization, the transmission data end flag is set to 1, the external synchronization trigger input of GLink is released, the GLink check for exchange end requirement flag is set to 1, the GLink transmission initialization requirement is cleared to 0, and the set message count within the GLink exchange group is cleared. GLink starts the next check timer after checking whether the exchange is completed. When it is detected that the exchange is completed, the enable of the GLink external synchronization trigger input is set to 0.
[0130] Figure 12 This is the timing diagram of data transmission and reception in the embodiments of the present disclosure. As Figure 12 shown, after the co-processor prefetches a message, parses and executes it, and then prefetches the next message. When the last message in the exchange group is parsed, the exchange group transmission initialization is started and the synchronization trigger input is enabled; it periodically queries whether the exchange group transmission is completed. At the same time, the exchange result of the previous exchange group is read into the secondary cache. When the communication results of an exchange group are all read into the secondary cache, the cache manager transfers the communication results of this exchange group from the secondary cache to the primary cache. When it is detected that the exchange group transmission is completed, the co-processor then prefetches the next group of messages.
[0131] In Figure 12 it can be seen that after the co-processor sets the data to be transmitted and starts the transmission initialization, during the period when the equal-length communication data is completed, the result of the previous data communication is read back into the secondary cache. Compared with receiving data after waiting for the transmission to complete, the utilization rate during the time slice of waiting for the transmission to complete can be improved.
[0132] Through the above solution, when the readable messages in the first terminal node 171 are greater than the preset message number threshold, the communication result of the first terminal node 171 can be received. When there is a communication instruction in the first secondary buffer 161 in the downlink direction and the write request flag bit of the first terminal node 171 is in the existence request state, an instruction download request can be sent to the instruction processing module 1513. When the initialization request flag bit is in the existence request state, the transceiver process handling module 1512 can perform initialization settings on the first terminal node 171. When the exchange end request flag bit is in the existence request state and the timing has ended, it can be checked whether the first terminal node 171 has sent the communication result to the first secondary buffer 161.
[0133] In some embodiments, the cache controller 130 includes: a register module 131 and a message control module 132;
[0134] The register module 131 is configured to control the write pointer of the first-level buffer in the downlink direction and the read pointer of the first-level buffer in the uplink direction;
[0135] The message control module 132 is configured to control the read pointer of the first-level buffer in the downlink direction and the write pointer of the first-level buffer in the uplink direction;
[0136] The register module 131 is configured to control the read pointer of the secondary buffer in the downlink direction and the write pointer of the secondary buffer in the uplink direction;
[0137] The message control module 132 is configured to control the write pointer of the secondary buffer in the downlink direction and the read pointer of the secondary buffer in the uplink direction.
[0138] In specific implementation, all the first-level caches share the same storage area, and multiple areas are divided in a piece of DDR memory; each area corresponds to the downlink (H2C Host to Card) or uplink (C2H Card to Host) first-level cache of an accelerator respectively. The secondary cache is the on-chip memory BRAM, and the downlink (H2C) and uplink (C2H) secondary caches of the same accelerator share a piece of BRAM.
[0139] Figure 13 It is a schematic structural diagram of the cache controller according to an embodiment of the present disclosure. As Figure 13 shown, the write pointer of the first-level buffer in the downlink direction (H2C direction) is controlled by the host computer through a register, that is, when the host computer writes data into the first-level buffer through DMA, the write pointer of the first-level buffer in the H2C direction is updated; the read pointer of the first-level buffer in the H2C direction is controlled by the message control module in the cache management IP, that is, when data is read from the first-level buffer to the corresponding secondary buffer, the message control module updates the read pointer of the first-level buffer.
[0140] Similarly, the write pointer of the first-level cache in the C2H direction is controlled by the message control module in the cache management IP, and the read pointer is controlled by the host computer. The write pointer of the second-level cache in the H2C direction and the read pointer of the second-level cache in the C2H direction are controlled by the message control module in the cache management IP. The write pointer of the second-level cache in the C2H direction and the read pointer of the second-level cache in the H2C direction are controlled by the coprocessor.
[0141] Through the above solution, the register module 131 can control the write pointer of the first-level buffer in the downlink direction and the read pointer of the first-level buffer in the uplink direction. The message control module 132 can control the read pointer of the first-level buffer in the downlink direction and the write pointer of the first-level buffer in the uplink direction. The register module 131 can control the read pointer of the second-level buffer in the downlink direction and the write pointer of the second-level buffer in the uplink direction. The message control module 132 can control the write pointer of the second-level buffer in the downlink direction and the read pointer of the second-level buffer in the uplink direction.
[0142] In some embodiments, the cache controller 130 is further configured to:
[0143] Judge the cache access request;
[0144] In response to determining that the cache access request is in the uplink direction, read and parse the configuration instruction in the message instruction of the second-level buffer in the uplink direction, write the current timestamp and the current state into the first-level buffer in the uplink direction, read the communication instruction in the message instruction, and write the communication instruction into the first-level buffer in the uplink direction;
[0145] In response to determining that the cache access request is in the downlink direction, read and parse the configuration instruction in the message instruction of the first-level buffer in the downlink direction, write the current timestamp and the current state into the second-level buffer in the downlink direction, read the communication instruction in the message instruction, and write the communication instruction into the second-level buffer in the downlink direction.
[0146] During specific implementation, the register module 131 for cache management mainly configures the base addresses and sizes of each cache, including the base addresses and sizes of each level-1 cache in the H2C direction, the base addresses and sizes of the level-1 cache in the C2H direction, the base addresses and sizes of each level-2 cache in the H2C direction, the base addresses and sizes of the level-2 cache in the C2H direction, and the host computer can set the write pointer of the level-1 cache in the H2C direction and the read pointer of the level-1 cache in the C2H direction in the cache management register module. Configure the cache management module to be enabled; configure each coprocessor to be enabled; configure the minimum number of messages to be moved from the level-2 cache to the level-1 cache. The host computer can obtain the following information by reading the corresponding registers in the cache management register module: the current read pointer position of the level-1 cache in the H2C direction, the current write pointer position of the level-1 cache in the C2H direction, the read and write pointer positions of each level-2 cache in the H2C and C2H directions; the number of readable messages and writable messages in each current cache; and status information such as whether there is an error status and the error count during the data processing process.
[0147] The level-1 cache control and level-2 cache control module calculates the number of readable messages and writable messages in the current cache, as well as the empty flag and full flag of the current cache according to the input cache read pointer and write pointer.
[0148] Figure 14 It is a schematic structural diagram of the message control module according to an embodiment of the present disclosure. As Figure 14 shown, the cache access process control includes: calculating the read pointer address of the level-1 cache in the H2C direction, calculating the write pointer address of the level-1 cache in the C2H direction, calculating the write pointer address of the level-2 cache 1 in the H2C direction, calculating the write pointer address of the level-2 cache 2 in the H2C direction,..., calculating the write pointer address of the level-2 cache n in the H2C direction, calculating the read pointer address of the level-2 cache 1 in the C2H direction, calculating the read pointer address of the level-2 cache 2 in the C2H direction,..., calculating the read pointer address of the level-2 cache n in the C2H direction. The cache access process control also includes: high-priority cache access round-robin arbitration, medium-priority cache access round-robin arbitration, judging the read demand of the level-2 cache 1 in the C2H direction, judging the read demand of the level-2 cache 2 in the C2H direction,..., judging the read demand of the level-2 cache n in the C2H direction, judging the write demand of the level-2 cache 1 in the H2C direction, judging the write demand of the level-2 cache 2 in the H2C direction,..., judging the write demand of the level-2 cache n in the H2C direction.
[0149] Each level-2 cache control read and write demand judgment will give the access priority according to the proportion of the current readable / writable message count. For the level-2 cache in the H2C direction, when the number of its writable messages is above the preset value (i.e., when it is half-empty), its access demand priority is high; when the number of its writable messages is below the preset value, its access priority is normal; when its writable messages are full, it has no access demand.
[0150] For the secondary cache in the C2H direction, when the number of readable messages is above the preset value (almost full), its access priority is the highest; when the number of readable messages is below the preset value, its access priority is normal; when the readable area is empty, there is no access requirement.
[0151] Figure 15 The flowchart shows the data processing of the message control module according to the embodiments of the present disclosure. As Figure 15 shown, it is determined whether the message control module has a high-priority request. If there is a high-priority request, according to the current high-priority request chip selection, an access request is initiated, and it is determined whether the access ends; when the access ends, the read or write pointer of the relevant cache is updated, and the high-priority chip selection is updated. If there is no high-priority request, it is determined whether there is an access request; when there is an access request, according to the current normal-level request chip selection, an access request is initiated, and it is determined whether the access ends; when the access ends, the read or write pointer of the relevant cache is updated, and the normal-level chip selection is updated.
[0152] According to the access priorities of each secondary cache, a round-robin access is first performed on the high-priority secondary cache. When there is no secondary cache with priority access, a round-robin access is performed on the medium-priority secondary cache.
[0153] Figure 16 The flowchart shows the data reading and writing of the message reading and writing module according to the embodiments of the present disclosure. As Figure 16 shown, after receiving an access request, the message reading and writing module for cache management determines the direction of the access request and obtains the corresponding cache read address pointer. For the cache access request in the H2C direction, first, the configuration instruction information in the message instruction is read from the H2C-direction primary cache, and the secondary cache corresponding to the target coprocessor of this message instruction, as well as information such as the length and data transfer direction of this message instruction, are obtained; secondly, the configuration instruction information and the current timestamp, etc., are written into the corresponding secondary cache together; finally, the data instruction in the primary cache is read and written into the secondary cache. For the cache access request in the C2H direction, first, the configuration instruction information in the message instruction is read from the corresponding C2H-direction secondary cache, and information such as the length and primary cache address of this message instruction is obtained by parsing the configuration instruction; the configuration instruction information, the current timestamp, and the status information, etc., are written into the corresponding C2H-direction primary cache; finally, the data instruction in the secondary cache is read and written into the primary cache.
[0154] Through the above solution, by judging the cache access request, it is possible to accurately identify whether the cache access request is in the upward direction or the downward direction. When the cache access request is in the upward direction, the communication instruction can be written into the primary cache in the upward direction. When the cache access request is in the downward direction, the communication instruction can be written into the secondary cache in the downward direction.
[0155] In some embodiments, the synchronization trigger module 140 includes: a clock synchronization module 141, a first trigger signal configuration module 142, and a second trigger signal configuration module 143;
[0156] The clock synchronization module 141 is specifically configured to:
[0157] Determine whether a periodic synchronization signal input from the outside is received;
[0158] In response to determining that a periodic synchronization signal input from the outside is received, adjust an initial periodic pulse signal according to the periodic synchronization signal to obtain a target periodic pulse signal;
[0159] Send the target periodic pulse signal to the first trigger signal configuration module 142 and the second trigger signal configuration module 143, so that the first trigger signal configuration module 142 and the second trigger signal configuration module 143 synchronously control the first terminal node 171 and the second terminal node 172 to perform synchronous communication.
[0160] Specifically, the periodic trigger signal sent by GLink is generated by hardware logic. The synchronization trigger source has a synchronization extension function, can receive external periodic signal input or output periodic synchronization signals externally, and realizes the cascade function of synchronous signals of multiple modules, so as to realize the synchronous operation of more coprocessors. Figure 17 This is a schematic structural diagram of the synchronization trigger module according to an embodiment of the present disclosure. As Figure 17 shown, the synchronization trigger module 140 includes: a clock synchronization module 141, a trigger signal configuration module 1, a trigger signal configuration module 2, a trigger signal configuration module 3, and a trigger signal configuration module 4.
[0161] The clock synchronization module can realize the input and output of a 1ms periodic synchronization pulse signal and generate a 1us periodic pulse locally. Figure 18 This is a flowchart of the clock synchronization module according to an embodiment of the present disclosure for synchronous communication. As Figure 18As shown, first, configure the selection of the synchronization signal source. An open-drain signal or an independent input / output signal can be selected. The open-drain signal is mainly for synchronous interconnection between different boards within the same PXIe chassis. The independent input / output signal is mainly for synchronous signal interconnection between different chassis or at a relatively long distance. Secondly, configure whether the synchronization signal source is generated internally by the board or input externally. For the synchronization signal source generated internally by the board, wait for the host computer to configure the timestamp to be cleared. After receiving the cleared timestamp, clear the local timestamp and output a timestamp clear pulse signal externally. Then the local timer starts to work, outputs a periodic synchronization pulse signal with a period of 1 ms externally, and after waiting for the configured delay, outputs a synchronization pulse signal with a period of 1 ms and a pulse timing signal with a period of 1 us within the board. When the synchronization signal source is input externally, wait for the externally input timestamp clear pulse signal, wait for the first externally input synchronization pulse signal. After the externally input synchronization pulse signal passes through the configured delay, start generating a local synchronization pulse signal with a period of 1 ms and a pulse timing signal with a period of 1 us. At the same time, dynamically adjust the period of the locally output 1 us periodic pulse counting signal according to the externally input 1 ms periodic synchronization pulse signal.
[0162] Figure 19 This is the flowchart of dynamically adjusting the local pulse period in the embodiment of the present disclosure. As Figure 19 shown, adjust the timing period of the local 1 us pulse to achieve acceleration or deceleration of the local timing clock. The clock difference between two boards will not exceed 1000 ppm (worst case). The local clock period is 8 ns, the local 1 us timing is 125 clock cycles, and the local 1 ms timing is 125000 clock cycles. When the 1 ms periodic synchronization pulse is input from an external board, in the worst case, the difference in the number of local clock cycles will not exceed 125. Therefore, the speed of the local timing clock can be adjusted by adjusting the number of clocks of the first m 1 us pulses. 1000 ppm * 125000 = 125. The specific method is to first determine the number of timing clock cycles of the externally input 1 ms synchronization pulse signal under the local clock, measure the number of clock cycles m that differ from 1 ms under the local clock. When the local clock is fast, the number of clock cycles of the first m subsequent output 1 us pulses is increased by 1, that is, the local clock cycle number of the first m 1 us pulses becomes 126, thereby slowing down the local timing clock. If the local clock is slow, the number of clock cycles of the first m subsequent output 1 us pulses is decreased by 1, that is, the local clock cycle number of the first m 1 us pulses becomes 124, thereby accelerating the local 1 timing clock.
[0163] The trigger signal configuration module can generate periodic or aperiodic trigger signals based on the 1 us pulse input signal, and the delay time of the trigger signal relative to the synchronization pulse is adjustable, so as to generate the synchronization trigger signals required for each coprocessor. Figure 20 This is the schematic diagram of the trigger signal in the embodiment of the present disclosure. InFigure 20 The figure shows a schematic diagram of the hardware synchronization clock transmission of 4 coprocessors under the synchronous trigger signals of different cycles. The transmission cycle of coprocessor 1 is 250 us, the transmission cycles of coprocessor 2 and coprocessor 4 are 1 ms, and the transmission cycle of coprocessor 3 is 100 ms. The output trigger signal cycles of the synchronization trigger modules 1 to 4 are set to 250 us, 1 ms, 100 ms, and 1 ms respectively; after the first periodic clock synchronization pulse signal after the timestamp clearing signal, the corresponding synchronous trigger signals are output after a delay.
[0164] The messages sent and received by the coprocessor are all transmitted in the form of instructions. The instructions are mainly divided into three types: instructions with 1 operand, instructions with 2 operands, and instructions with 3 operands. The instruction code occupies 8 bits and is located in the highest 8 bits of each instruction, that is, bit24 to bit31; when there is 1 operand, the operand length is 24 bits; it is located in bit0 to bit23 of the instruction; when there are 2 operands, the operand lengths are each 12 bits; they are located in bit12 to bit23 and bit0 to bit11 of the instruction respectively; when there are 3 operands, the operand lengths are each 8 bits; they are located in bit16 to bit23; bit8 to bit15; bit0 to bit7 of the instruction respectively. The instruction code controls the cache manager and the coprocessor to perform corresponding operations and read back the corresponding intermediate key status information and data results.
[0165] Multiple instructions that control the same GLink communication are combined together to form a message MSG, and multiple messages MSG that need to be sequentially sent after one trigger are combined together to form a message group Closure. Within a message group, the sequence number of the INIT_Set_Target in the message group increases from 1 to the total number of messages in the message group.
[0166] For the convenience of coprocessor processing, the sent message has a fixed length of 512 bytes. The received message has a length of 1024 bytes.
[0167] The sent message is shown in Table 1 below. The first 4 are configuration instructions, and their order can be exchanged. The 5th to 8th are placeholder instructions, and the data after the 9th is the data to be sent. If the data is insufficient, it is occupied by placeholder instructions; their order cannot be changed arbitrarily.
[0168] Table 1 Instruction Configuration Table of Sent Messages
[0169]
[0170] The sending instruction group is configured by the host computer. To send an instruction, it is necessary to configure the target coprocessor of the message, the sequence number within the message group, the number of messages in the message group; configure the transmission direction, transmission sub-address, transmission length, etc. of the GLink communication message; configure the target site ID number of the GLink communication message; fill in the corresponding data to be sent according to the transmission direction and length of the communication message.
[0171] Table 2 Receiving Message Instruction Configuration Table
[0172]
[0173]
[0174] The cache manager and the coprocessor each process one message at a time. After reading the message, obtain its configuration information and perform corresponding operations according to its configuration information.
[0175] Multiple instruction group Closures include: GLK0_Closure1, GLK0_Closure2, GLK0_Closure3, ……, GLK0_Closure m. Each GLK0_Closure includes: MSG1, MSG2, MSG3, and MSGn.
[0176] A sending message queue can be formed for each GLink terminal node. After sending the messages of each GLink terminal node to the corresponding area of the first-level cache of the board, the cache controller can be enabled to start working. Figure 21 Flowchart of forming a sending message queue for the host computer in the embodiment of the present disclosure. As Figure 21 shown, when there is a message to be sent, apply for a continuous sending buffer; form a sending message structure; add the structure to the sending queue and fill it into the buffer; determine whether all the sending messages have been set or the applied buffer is full; if so, set the DMA sending address and length according to the current read / write address of the first-level cache of the target GLink terminal node, and start the DMA sending; if not, re-form the sending message structure.
[0177] Figure 22 Schematic diagram of the timestamp for the coprocessor to control GLink communication in the embodiment of the present disclosure. Figure 22 Timestamp situations for 4 coprocessors to control GLink communication; each message group corresponding to each GLink node is a group of 20 messages, and the synchronization trigger signal period corresponding to each coprocessor is 1000us. Figure 22The horizontal axis represents the messages sent, that is, each column represents the timestamps of the same message. The vertical axis represents time, with the unit of μs, and the timestamps are normalized to start from 0. The GLink communication is configured to wait for a synchronous external trigger after every 20 messages at each site. Each GLink site sends 20 messages each time, and sends 3 times. A total of 240 messages are sent by 4 GLink sites. In Figure 22 it respectively shows (from bottom to top in turn) the actual timestamps of the transfer from the secondary cache to GLink, the start and end timestamps of each message in GLink communication, the actual timestamps of the transfer of the GLink communication result to the secondary cache, and the timestamps of the transfer from the secondary cache to the primary cache.
[0178] Through multiple coprocessors, parallel transceiver operations of multiple GLink terminal nodes are realized, so that the operations of each GLink terminal node no longer wait for the operations of other terminal nodes to complete, improving the data transceiver efficiency. The CPU can interact with the board for instructions, data, and results through a more efficient DMA (Direct Memory Access) data transfer method, improving the access efficiency of PCIe and reducing the strong real-time requirements for the CPU. Through the processing of the trigger signal by the synchronization trigger IP and the acceleration control IP, synchronous data transceiver of multiple GLink terminal nodes of multiple boards can be realized.
[0179] Figure 23 It is a schematic structural diagram of the board of the embodiment of the present disclosure. As Figure 23 shown, the board includes: a PXIe module, an FPGA module, an optical-electric conversion module, a power supply module, a storage module, and a clock module. Among them, in the PXIe bus module, the PXIe bus is the system bus of the board, used for the host computer to control and exchange data, and provides backplane PXI trigger, and provides backplane power supply for the board. The FPGA module integrates a PCIe bus IP core, a memory control IP core, a GLink communication IP core, a cache management IP core, an acceleration control IP core, a synchronization trigger IP core, etc. The power supply module supplies power to the PXIe bus and consists of DC / DC modules to form digital power supplies such as +1.0V, +1.2V, +1.5V, +1.8V, +3.3V, etc. The storage module consists of multiple DDR3 chips, providing a large-capacity cache for data acceleration of the board. The clock module provides a reference clock for the operation of the FPGA and the memory of the board. The optical-electric conversion module provides an optical interface for external communication of GLink nodes.
[0180] Through the above solution, by judging the received periodic synchronization signal, it is possible to accurately judge whether the periodic synchronization signal is externally input or internally generated. When receiving an externally input periodic synchronization signal, the initial periodic pulse signal is adjusted according to the periodic synchronization signal to obtain a target periodic pulse signal, making the target periodic pulse signal more accurate.
[0181] Through the above embodiments, the first coprocessor and the second coprocessor can simultaneously configure the corresponding terminal nodes, thereby shortening the instruction sending time. The first coprocessor and the second coprocessor can simultaneously read the communication results from the corresponding terminal nodes, thereby shortening the instruction reading time, and there is no need to wait for the writing and reading of other terminal nodes.
[0182] For the convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0183] Based on the same inventive concept, corresponding to the device in any of the above embodiments, the present disclosure also provides a synchronous communication method based on multiple coprocessors.
[0184] Reference Figure 24 , the synchronous communication method based on multiple coprocessors is applied to the synchronous communication device based on multiple coprocessors described in the above embodiments; the device includes: a main processor, a primary buffer, a cache controller, a synchronization trigger module, a first coprocessor, a first secondary buffer, a first terminal node, a second coprocessor, a second secondary buffer, and a second terminal node; the method includes:
[0185] Step 201, the main processor sends a communication instruction to the primary buffer;
[0186] Step 202, the cache controller distributes the communication instruction from the primary buffer to the first secondary buffer and the second secondary buffer;
[0187] Step 203, the first coprocessor reads the communication instruction from the first secondary buffer and configures the communication instruction to the first terminal node;
[0188] Step 204, the second coprocessor reads the communication instruction from the second secondary buffer and configures the communication instruction to the second terminal node;
[0189] Step 205, the synchronization trigger module controls the first terminal node and the second terminal node to perform synchronous communication according to the communication instruction to obtain a communication result;
[0190] Step 206, the first coprocessor reads the communication result from the first terminal node and stores the communication result in the first secondary buffer;
[0191] Step 207, the second coprocessor reads the communication result from the second terminal node and stores the communication result in the second secondary buffer;
[0192] Step 208: The cache manager collects the communication results in the first secondary cache and the second secondary cache into the primary cache;
[0193] Step 209: The main processor collects the communication results in the primary cache into the host computer, and uses the host computer to parse and process the communication results to obtain a parsing result.
[0194] In specific implementation, as Figure 5 shown, when the GLink communication includes GLink1, GLink2, and GLink3, the data communication method includes: The main processor packs the communication instruction data to be sent and sends it to the primary cache through direct memory access (DMA for short); the cache manager distributes the communication instruction data in the primary cache to each secondary cache; the first coprocessor reads the communication instruction data from the secondary cache 1 and configures the GLink1 communication data; the second coprocessor reads the communication instruction data from the secondary cache 2 and configures the GLink2 communication data; the third coprocessor reads the communication instruction data from the secondary cache 3 and configures the GLink3 communication data; start all GLink transmissions; the first coprocessor reads the GLink1 communication result and stores it in the secondary cache 1; the second coprocessor reads the GLink2 communication result and stores it in the secondary cache 2; the third coprocessor reads the GLink3 communication result and stores it in the secondary cache 3; the cache manager receives the communication results in each secondary cache into the primary cache; the main processor receives the communication results in the primary cache into the host computer through DMA and parses them.
[0195] Through the above embodiments, the first coprocessor and the second coprocessor can configure the corresponding terminal nodes simultaneously, thereby shortening the instruction sending time. The first coprocessor and the second coprocessor can simultaneously read the communication results from the corresponding terminal nodes, thereby shortening the instruction reading time, without waiting for the writing and reading of other terminal nodes.
[0196] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario, and multiple devices cooperate with each other to complete it. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0197] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0198] The method of the above embodiments is applied to the corresponding synchronous communication device based on multiple coprocessors in any of the foregoing embodiments, and has the beneficial effects of the corresponding device embodiments, which will not be elaborated here.
[0199] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the synchronous communication method based on multiple coprocessors described in any of the above embodiments when executing the program.
[0200] Figure 25 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0201] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0202] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0203] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure), or can be externally connected to the device to provide corresponding functions. Among them, the input devices can include keyboards, mice, touchscreens, microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0204] The communication interface 1040 is used to connect to the communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB (Universal Serial Bus), network cable, etc.), or can achieve communication through wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0205] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0206] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0207] The electronic device of the above embodiment is used to implement the corresponding synchronous communication method based on multiple coprocessors in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0208] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the synchronous communication method based on multiple coprocessors as described in any of the foregoing embodiments.
[0209] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0210] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the synchronous communication method based on multiple coprocessors described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0211] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a computer program product, including computer program instructions, which when run on a computer, cause the computer to execute the synchronous communication method based on multiple coprocessors described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0212] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0213] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be executed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0214] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0215] It should be understood that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0216] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, and they are not provided in detail for the sake of brevity.
[0217] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation manners of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0218] Although the present disclosure has been described in connection with specific embodiments of the present disclosure, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) can be used with the embodiments discussed.
[0219] The embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the present disclosure. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A synchronous communication device based on multiple coprocessors, characterized in that The device includes: a main processor, a first-level buffer, a cache controller, a synchronization trigger module, a first coprocessor, a first second-level buffer, a first terminal node, a second coprocessor, a second second-level buffer, and a second terminal node; The main processor is configured to send communication instructions to the first-level buffer; The cache controller is configured to distribute the communication instructions from the first-level buffer to the first second-level buffer and the second second-level buffer; The first coprocessor is configured to read the communication instructions from the first second-level buffer and configure the communication instructions into the first terminal node; The second coprocessor is configured to read the communication instructions from the second second-level buffer and configure the communication instructions into the second terminal node; The synchronization trigger module is configured to control the first terminal node and the second terminal node to perform synchronous communication according to the communication instructions to obtain a communication result; The first coprocessor is configured to read the communication result from the first terminal node and store the communication result in the first second-level buffer; The second coprocessor is configured to read the communication result from the second terminal node and store the communication result in the second second-level buffer; The cache manager is configured to collect the communication results in the first second-level buffer and the second second-level buffer into the first-level buffer; The main processor is configured to collect the communication results in the first-level buffer into the host computer, and use the host computer to parse and process the communication results to obtain a parsing result.
2. The device according to claim 1, characterized in that, The first coprocessor includes: a register module, a transceiver process handling module, and an instruction processing module; The instruction processing module is configured to read the communication instructions from the first second-level buffer, and parse and process the communication instructions to obtain an instruction configuration position; The register module is configured to configure the communication instructions to the instruction configuration position in the first terminal node, and record the read address of the second-level buffer in the downlink direction and the write address of the second-level buffer in the uplink direction of the first coprocessor; The instruction processing module is configured to determine that the number of messages in the exchange group of the communication instructions is equal to the message sequence number, and send a configuration completion notification to the transceiver process handling module; The transceiver process handling module is configured to perform initialization settings on the first terminal node, so that the first terminal node performs synchronous communication under the control of the synchronization trigger module; The instruction processing module is configured to read the communication results from the first terminal node in a parallel manner, and store the communication results in the first second-level buffer in the uplink direction.
3. The device according to claim 2, wherein The instruction processing module is further configured to: receive a request to be processed, and perform judgment processing on the request to be processed; In response to determining that the request to be processed is an instruction sending request, read a communication instruction from the first-level secondary buffer in the downlink direction, write a configuration instruction timestamp to the first-level secondary buffer in the uplink direction, and configure the communication instruction to the data stack of the first terminal node; In response to determining that the instruction to be processed is a result reading instruction, read a message instruction group from the first-level secondary buffer in the uplink direction, write a timestamp and the current status to the first-level secondary buffer in the uplink direction, and read the communication result from the stack of the first terminal node; In response to determining that the instruction to be processed is an exchange end instruction, check the stack exchange count of the first terminal node; In response to determining that the instruction to be processed is an initialization instruction, configure the stack pointer and stack count of the first terminal node.
4. The device according to claim 2, characterized in that, The transceiver process handling module is further configured to: In response to determining that the readable messages in the first terminal node are greater than a preset message count threshold, receive the communication result of the first terminal node; In response to determining that there is a communication instruction in the first-level secondary buffer in the downlink direction and the write request flag bit of the first terminal node is in the existence request state, send an instruction distribution request to the instruction processing module, so that the instruction processing module configures a communication instruction to the first terminal node according to the instruction distribution request; In response to determining that the initialization request flag bit is in the existence request state, the transceiver process handling module performs an initialization setting on the first terminal node; In response to determining that the exchange end request flag bit is in the existence request state and the timing has ended, check whether the first terminal node has sent the communication result to the first-level secondary buffer.
5. The device according to claim 1, characterized in that, The cache controller includes: a register module and a message control module; The register module is configured to control the write pointer of the first-level buffer in the downlink direction and the read pointer of the first-level buffer in the uplink direction; The message control module is configured to control the read pointer of the first-level buffer in the downlink direction and the write pointer of the first-level buffer in the uplink direction; The register module is configured to control the read pointer of the second-level buffer in the downlink direction and the write pointer of the second-level buffer in the uplink direction; The message control module is configured to control the write pointer of the second-level buffer in the downlink direction and the read pointer of the second-level buffer in the uplink direction.
6. The device according to claim 1, characterized in that The cache controller is further configured to: Judge a cache access request; In response to determining that the cache access request is in the uplink direction, read and parse the configuration instruction in the message instruction of the second-level buffer in the uplink direction, write the current timestamp and the current status to the first-level buffer in the uplink direction, read the communication instruction in the message instruction, and write the communication instruction to the first-level buffer in the uplink direction; In response to determining that the cache access request is in the downlink direction, read and parse the configuration instruction in the message instruction of the first-level buffer in the downlink direction, write the current timestamp and the current status to the second-level buffer in the downlink direction, read the communication instruction in the message instruction, and write the communication instruction to the second-level buffer in the downlink direction.
7. The device according to claim 1, characterized in that, The synchronization trigger module includes: a clock synchronization module, a first trigger signal configuration module, and a second trigger signal configuration module; The clock synchronization module is specifically configured to: Determine whether a periodic synchronization signal is received from an external input; In response to determining that a periodic synchronization signal is received from an external input, adjust an initial periodic pulse signal according to the periodic synchronization signal to obtain a target periodic pulse signal; Send the target periodic pulse signal to the first trigger signal configuration module and the second trigger signal configuration module, so that the first trigger signal configuration module and the second trigger signal configuration module synchronously control the first terminal node and the second terminal node to perform synchronous communication.
8. A synchronous communication method based on multiple coprocessors, characterized in that Applied to the synchronous communication device based on multiple coprocessors as claimed in claims 1 to 7; the device includes: a main processor, a first-level buffer, a cache controller, a synchronization trigger module, a first coprocessor, a first second-level buffer, a first terminal node, a second coprocessor, a second second-level buffer, and a second terminal node; the method includes: The main processor sends a communication instruction to the first-level buffer; The cache controller distributes the communication instruction from the first-level buffer to the first second-level buffer and the second second-level buffer; The first coprocessor reads the communication instruction from the first second-level buffer and configures the communication instruction into the first terminal node; The second coprocessor reads the communication instruction from the second second-level buffer and configures the communication instruction into the second terminal node; The synchronization trigger module controls the first terminal node and the second terminal node to perform synchronous communication according to the communication instruction to obtain a communication result; The first coprocessor reads the communication result from the first terminal node and stores the communication result in the first second-level buffer; The second coprocessor reads the communication result from the second terminal node and stores the communication result in the second second-level buffer; The cache manager collects the communication results in the first second-level buffer and the second second-level buffer into the first-level buffer; The main processor collects the communication results in the first-level buffer into a host computer, and uses the host computer to parse and process the communication results to obtain a parsing result.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the method as claimed in claim 8.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method as claimed in claim 8.
Citation Information
Cited By
Storage area resource allocation method and system based on Glink CtlrNT mode
CN121434118A