Time-based synchronization descriptors
By including timing data and operators in the work request entry, peripheral devices such as NICs and accelerators execute the work request based on the timing data, solving the problem of high-precision time synchronization between endpoints in high-performance networks and realizing flexible and efficient TDMA processing.
Patent Information
- Application Number
- CN202310034356.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-02-09
- Filing Date
- 2023-01-10
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-01-10
AI Technical Summary
Existing technologies struggle to achieve high-precision time synchronization between endpoints in high-performance networks, resulting in time division multiplexing (TDMA) processing being cumbersome and inflexible, unable to create timestamps for arbitrary times, and incurring high processing costs.
By including timing data and operators in the job request entries, peripheral devices such as NICs and accelerators execute job requests based on the timing data, achieving precise time synchronization and workload processing in conjunction with hardware clocks and processing circuits.
It achieves high-precision time synchronization between endpoints, improves the flexibility and efficiency of TDMA processing, and reduces processing costs.
Smart Images

Figure CN116582595B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computer systems, and particularly, but not exclusively, to the timed execution of workloads. Background Technology
[0002] Various techniques for implementing Time Division Multiplexing (TDM) protocols in networks such as Ethernet are known in the art. For example, “Practical TDMA for Data Center Ethernet” by Vattikonda et al., published in April 2012 by the Department of Computer Science and Engineering at the University of California, San Diego, describes the design and implementation of a TDMA Media Access Control (MAC) layer for commercial Ethernet hardware, which allows end hosts to be exempt from TCP reliability and congestion control.
[0003] In another example, U.S. Patent Application Publication 2019 / 0319730 describes techniques for operating a Time Division Multiplexing (TDM) MAC module, including examples of facilitating the use of shared resources allocated to ports of a network interface based on a time slot mechanism, wherein the shared resources are allocated to packet data received or transmitted through the ports of the network interface.
[0004] U.S. Patent Publication 2021 / 0297151 by Levi et al. describes a network element comprising one or more network ports, a network timing circuit, and packet processing circuitry. The network ports are configured to communicate with a communication network. The network timing circuitry is configured to track a network time defined in the communication network. In some embodiments, the packet processing circuitry is configured to receive definitions of one or more time slots synchronized with the network time and to send outbound packets to the communication network according to the time slots. In some embodiments, the packet processing circuitry is configured to process inbound packets received from the communication network according to the time slots. Summary of the Invention
[0005] According to embodiments of this disclosure, a system including a peripheral device is provided, the peripheral device including a hardware clock and processing circuitry, the processing circuitry being configured to: read a given work request entry stored in a plurality of work request entries in at least one work queue in memory, the given work request entry including timing data and an operator, the timing data indicating the time when the work request should be executed; retrieve a clock value from the hardware clock; and execute the work request having a workload when the execution of the work request is timed in response to the timing data, the operator, and the retrieved clock value.
[0006] Furthermore, according to embodiments of this disclosure, the system includes a plurality of host devices, each host device including: an interface for connecting to and transmitting data with the peripheral device; and a processor for running at least one software application configured to write host-specific work request entries from the work request entries into at least one corresponding work queue in the at least one work queue.
[0007] Further according to embodiments of this disclosure, the system includes a host device comprising: an interface for connecting to and transmitting data with the peripheral device; and a processor for running at least one software application configured to write the job request entries into the at least one job queue in the memory, wherein the peripheral device includes an interface configured to connect to and transmit data with the host device.
[0008] Additionally, according to embodiments of this disclosure, the host device includes the memory.
[0009] Furthermore, according to embodiments of this disclosure, at least one software application is configured to write the timing data in Coordinated Universal Time (UTC) format.
[0010] Furthermore, according to embodiments of this disclosure, the at least one software application is configured to generate corresponding timing data for writing to the corresponding entries in the work request entries of the at least one work queue in the memory in response to time-division multiplexing scheduling.
[0011] Furthermore, according to embodiments of this disclosure, the processing circuitry is configured to process corresponding timing data in response to a corresponding work request entry in the work request entry and a corresponding clock value retrieved from the hardware clock, and packets to be transmitted over the network, timed according to the time-division multiplexing schedule.
[0012] Furthermore, according to embodiments of this disclosure, the operator is selected from any one or more of the following: conditional operator; comparison operator; less than; greater than; equal to; including pattern; including periodic pattern; within a range; conforming mask.
[0013] Furthermore, according to embodiments of this disclosure, the mask includes masking the most significant bit of the retrieved clock value.
[0014] Furthermore, according to embodiments of this disclosure, the corresponding timing data of a corresponding work request entry in a work request entry defines the execution timing of the work requests included in other work request entries of the work request entry.
[0015] Furthermore, according to embodiments of this disclosure, the corresponding entry in the work request entry is a fence entry with no work request.
[0016] Furthermore, according to embodiments of this disclosure, the corresponding work request entry in the work request entry defines the execution timing of the corresponding work request included in the corresponding work request entry in the work request entry.
[0017] Furthermore, according to embodiments of this disclosure, at least some of the work request entries include references to locations in memory where the corresponding workload should be stored or distributed.
[0018] Furthermore, according to embodiments of this disclosure, the peripheral device includes a network interface controller and a network interface for sharing packets with remote devices via a packet data network.
[0019] Furthermore, according to embodiments of this disclosure, the peripheral device includes a data processing unit (DPU) that includes a network interface controller, a network interface, and at least a processing core.
[0020] Furthermore, according to embodiments of this disclosure, the network interface is configured to receive packets over the network, and the processing circuitry is configured to process the received packets in response to corresponding timing data and corresponding operators of the corresponding work request entry in the work request entry.
[0021] Additionally, according to one embodiment of this disclosure, the processing circuitry is configured to process packets to be transmitted over the network in response to corresponding timing data and corresponding operators of a corresponding work request entry in the work request entry, and the network interface is configured to transmit packets over the network.
[0022] Furthermore, according to embodiments of this disclosure, the processing circuitry includes a clock controller for training a hardware clock in response to a Precision Time Protocol (PTP).
[0023] Furthermore, according to embodiments of this disclosure, the processing circuitry includes a hardware accelerator for executing at least one of the job requests.
[0024] Furthermore, according to embodiments of this disclosure, the hardware accelerator is configured to perform any one or more of the following: encrypting the workload, decrypting the workload, calculating the cyclic redundancy check (CRC) of the workload, or calculating the cryptographic hash of the workload.
[0025] According to another embodiment of this disclosure, a system is also provided, comprising: a plurality of peripheral devices for interconnecting with each other via a network, and including respective hardware clocks to be synchronized with each other via the network; and a plurality of host devices connected to the network via the peripheral devices, the host devices including respective processors for running respective software applications to share data with each other via the network, each of the respective software applications being configured to write a respective job request entry to a respective job queue, the respective job request entry including a respective operator and respective timing data, the timing data indicating the time when the respective job request should be executed, wherein each of the peripheral devices includes processing circuitry for: reading the respective job request entry from the respective job queue; retrieving a respective clock value from the respective hardware clocks; and executing the respective job request with a workload, timed in response to the respective timing data, the respective operator, and the retrieved respective clock value.
[0026] Furthermore, according to embodiments of this disclosure, the corresponding software applications are configured to share data with each other over a network and generate corresponding work request entries in response to time-division multiplexing scheduling.
[0027] According to another embodiment of this disclosure, a method is also provided, comprising: reading a given work request entry stored in a plurality of work request entries in at least one work queue in a memory, the given work request entry including timing data and an operator, the timing data indicating the time at which the work request should be executed; retrieving a clock value from a hardware clock; and executing the work request having a workload when the execution of the work request is timed in response to the timing data, the operator, and the retrieved clock value. Attached Figure Description
[0028] The invention will be understood from the following detailed description, taken in conjunction with the accompanying drawings, in which:
[0029] Figure 1 This is a block diagram of a time synchronization computer system constructed and operated according to an embodiment of the present invention;
[0030] Figure 2 yes Figure 1 A block diagram of the host and peripheral devices in the system;
[0031] Figure 3 It is used for Figure 1 A schematic diagram illustrating time-division multiplexing scheduling in a system;
[0032] Figure 4 It includes Figure 2 A flowchart of the steps in the operation method of the host device;
[0033] Figure 5 It is used for Figure 1 A schematic diagram of the work queue in the system;
[0034] Figure 6 It is used for Figure 1 A diagram illustrating job request entries in the system;
[0035] Figure 7 It is used for Figure 1 A diagram illustrating job request fence entries in the system;
[0036] Figure 8 It includes Figure 2 Flowcharts of the steps in the operation method of peripheral devices; and
[0037] Figure 9 It is used for Figure 1 A schematic diagram of clock value masking in the system. Detailed Implementation
[0038] Overview
[0039] Communication networks such as the Enhanced Common Public Radio Interface (eCPRI), Optical Data Center Networks (ODCN), and IP video (e.g., SMPTE 2110) use Time Division Multiplexing (TDM) or sometimes Time Division Multiple Access (TDMA) for communication between endpoints, where multiple data sources share the same physical medium during different time intervals called time slots. Time Division Multiplexing can be used in various implementations. For example, according to the 5G standard, different queues can be scheduled for different time periods. For example, queues 1 and 2 in time period 1, queues 3 and 4 in time period 2, and so on.
[0040] TDMA multiplexing in high-performance networks requires good synchronization between endpoints, which is typically achieved through a high-precision time base. Dedicated circuits can also be used to transmit and receive data in TDM networks; however, such dedicated circuits can be expensive and inflexible.
[0041] One solution is to allow the network interface controller (NIC) to process packets at specific times set by an application running on the host device. U.S. Patent Publication 2021 / 0297151 by Levi et al., mentioned earlier, provides such a solution. However, because this solution is pace-based, it cannot create timestamps for any desired time, is computationally intensive for the hardware, and may not be scalable.
[0042] Embodiments of the present invention address at least some of the aforementioned problems by allowing timing data to be included in work request entries (stored in one or more work queues), enabling peripheral devices such as NICs, smart NICs, and / or accelerators to execute work requests based on timing data. Work requests may include requests to process packets to be sent, process packets to be received, encrypt and / or decrypt data, calculate CRC, and / or calculate cryptographic hashes.
[0043] In some embodiments, a software application running on a host device writes work request entries with timing data to one or more work queues, such as those stored in the host device's memory. A peripheral device reads the work request entries from the work queues and executes the work request based on the timing data. In some embodiments, a clock value from a local clock, such as a physical clock (e.g., a Precision Time Protocol (PTP) synchronized clock), can be retrieved and compared with a time value included in the timing data to determine whether the peripheral device should execute the work request. The clock and time values in the timing data can be in any suitable format, such as Coordinated Universal Time (UTC) format.
[0044] In some embodiments, operators are also written into job request entries that describe how the timing data should be applied. For example, the operator could be "greater than," indicating that the associated job request should be executed when the time is greater than the time indicated by the timing data. Operators can include any suitable operators, such as conditional operators, comparison operators, less than, greater than, equal to, including patterns (e.g., 100 microseconds per second), including periodic patterns, within a range (e.g., between X seconds and Y seconds), and masks, which will be described in more detail below.
[0045] In some embodiments, a timing value (and optionally an operator) may be included in a work request entry with a corresponding work request, such that for each work request entry, the timing value in the work request entry specifies the time at which the corresponding work request in that work request entry should be executed. Alternatively, or optionally, a work request entry may be a fence entry that includes timing data (and optionally an operator) without a corresponding work request (or a work request with a zero value). A fence entry specifies the execution timing (excluding timing data or operators) of a work request included in another work request entry. Fence entries can be used in time-division multiplexing, where different packets from different queues are processed in different time intervals.
[0046] In some embodiments, operators (e.g., in fence entries) may include masks. Masks can be applied to retrieved clock values. For example, if a mask specifies that all the most significant bits of a retrieved clock value, except for the millisecond value, are masked, the mask can be used to identify when the millisecond value is in the range of 0 to 100, such that the mask can be used to identify the time within the first 100 milliseconds per second. This can be used in a time-division multiplexing system where packets from queue A are transmitted in the first 100 milliseconds per second, while packets from queue B are transmitted in the last 100 milliseconds per second, and so on. Queue A may include fence entries specifying that the most significant bits are masked and checking the remaining bits to determine if the value is between 0 and 99. Queue B may include fence entries specifying that the most significant bits are masked and checking the remaining bits to determine if the value is between 100 and 199, and so on.
[0047] In some embodiments, multiple host devices may connect to a single peripheral device and write job request entries to one or more job queues for reading and execution by the peripheral device.
[0048] In some embodiments, multiple host devices may be interconnected on a network (e.g., via a network) via multiple peripheral devices having clocks synchronized with each other. The host devices may run their respective applications, which communicate with each other according to a schedule such as time-division multiplexing scheduling.
[0049] System Description
[0050] Now for reference Figure 1 This is a block diagram view of a time-synchronization computer system 10 constructed and operated according to an embodiment of the present invention. System 10 includes a plurality of host devices 12 and a plurality of peripheral devices 14. The peripheral devices 14 are interconnected via a network 16. Each peripheral device 14 may include a hardware clock 18. The hardware clocks 18 of the respective peripheral devices 14 are configured to synchronize with each other on the network 16 using a suitable time synchronization protocol such as PTP. The host devices 12 are configured to connect to the network 16 (and interconnect with each other) via the peripheral devices 14. The host devices 12 can respond to a time-division multiplexing schedule 20 described in more detail below (…). Figure 1 They run software applications to share data with each other over the network.
[0051] Now for reference Figure 2 , Figure 2 yes Figure 1The diagram shows one of the host devices 12 and the corresponding peripheral device 14 in system 10. The host device 12 includes a processor 22, memory 24, and an interface 26. The interface 26 (e.g., a peripheral bus interface) is configured to connect to and transmit data with the peripheral device 14. The processor 22 is configured to run at least one software application 28, which is configured to write (host-specific) job request entries to at least one corresponding job queue 30 in memory 24.
[0052] Peripheral device 14 includes processing circuitry 32, interface 34, hardware clock 18, and network interface 36. Interface 34 (e.g., a peripheral bus interface) is configured to connect to and transmit data with host device 12. Network interface 36 is configured to share packets with remote device 46 via a packet data network (e.g., network 16). Network interface 36 is configured to receive and transmit packets via network 16. Processing circuitry 32 may include one or more processing cores 38, hardware accelerator 40, network interface controller (NIC) 42, and clock controller 44. Hardware accelerator 40 is configured to execute reference... Figure 8 A more detailed description of the work request follows. The network interface controller 42 may include a physical layer (PHY) chip and a MAC chip (not shown) to process received packets and packets for transmission over network 16. The clock controller 44 is configured to train the hardware clock 18 in response to a clock synchronization protocol such as PTP. See reference... Figure 8 The function of the processing circuitry 32 is described in more detail. In some embodiments, the peripheral device 14 includes a data processing unit 48 (DPU) (e.g., a smart NIC), which includes the processing circuitry 32, an interface 34, a network interface 36, and a hardware clock 18.
[0053] In practice, some or all of these functions of the processing circuitry 32 may be combined in a single physical component, or alternatively, implemented using multiple physical components. These physical components may include hardwired or programmable devices, or a combination of both. In some embodiments, at least some functions of the processing circuitry 32 may be executed by a programmable processor under the control of suitable software. This software may be downloaded to the device electronically, for example, via a network. Alternatively or additionally, the software may be stored in a tangible, non-transitory computer-readable storage medium, such as optical, magnetic, or electronic memory.
[0054] Now for reference Figure 3 It is used for Figure 1 A schematic diagram of the time division multiplexing scheduling 20 in system 10. Figure 3This shows that packets from queues 1 and 2 are sent in time periods t1 and t3 (i.e., odd-numbered time periods), while packets from queues 3 and 4 are sent in time periods 2 (i.e., even-numbered time periods).
[0055] Now for reference Figure 4 , Figure 4 It includes Figure 2 The flowchart 50 shows the steps in the operation method of the host device 12.
[0056] Software application 28 running on processor 22 of host device 12 is configured to generate job request entries (box 52) as well as timing data and operators (box 54). The timing data can use any suitable time format. In some embodiments, software application 28 is configured to write the timing data according to UTC format. Software application 28 is configured to add timing data and operators to at least some job request entries (box 56). For example, timing data and corresponding operators can be added to every job request entry or only to some job request entries (e.g., fence entries), as referenced. Figure 5 For a more detailed description, software application 28 is configured to write work request entries into work queue 30 (box 58) in memory 24.
[0057] Work request entries can be written to different work queues 30 according to the time-division multiplexing schedule 20. For example, Figure 3 This illustrates how different queues can facilitate time-division multiplexing of packets or any suitable workload. In some embodiments, software application 28 is configured to generate corresponding timing data for writing appropriate work request entries to work queue 30 in response to time-division multiplexing schedule 20.
[0058] Operators can be selected from any suitable operators, such as one or more of the following: conditional operator; comparison operator; less than; greater than; equal to; including pattern; including periodic pattern; within a range; conformation mask. (See reference...) Figure 9 In more detail, the mask may include masking the most significant bit (and optionally some least significant bits) of the retrieved clock value.
[0059] In some embodiments, software application 28 sends a "doorbell" signal (e.g., writes to a specific memory location) to peripheral device 14, notifying peripheral device 14 that there is work to be performed on work queue 30. Peripheral device 14 then responds to the doorbell by reading entries from work queue 30 and performing the work according to timing data and operators, as referenced. Figure 8 More detailed description.
[0060] Therefore, in Figure 1In the context of multiple host devices 12, each host device 12 may include a corresponding processor 22 to run a corresponding software application 28 (i.e., each processor 22 of each host device 12 runs one of the software applications 28) to share data with each other via network 16. In some embodiments, the corresponding software applications 28 are configured to share data with each other via network 16 and in response to time-division multiplexing scheduling 20 ( Figure 1 This generates the corresponding job request entries.
[0061] In such Figure 1 In the illustrated multi-host environment, each corresponding software application 28 (run by the corresponding processor 22 of the corresponding host device 12) is configured to write a corresponding job request entry into a corresponding job queue 30. The job request entry includes a corresponding operator and corresponding timing data indicating the appropriate time when the corresponding job request should be executed. In the multi-host environment, each peripheral device 14 is configured to: read a job request entry (including the corresponding timing data and the corresponding operator) from the corresponding job queue 30 (i.e., the job queue 30 written for that peripheral device 14); retrieve a corresponding clock value from a corresponding hardware clock 18 (i.e., from the hardware clock 18 of that peripheral device 14); and execute the corresponding job request (i.e., the job request included in the read job request entry or the job request in another job request entry whose execution time is being timed by the read job request entry). The workload is timed in response to the corresponding timing data, the corresponding operator, and the retrieved corresponding clock value, as referenced. Figure 8 A more detailed description.
[0062] Now for reference Figure 5 It is used for Figure 1 A schematic diagram of work queue 30 in system 10.
[0063] Figure 5 Two work queues 30 are shown: work queue 30-1 and work queue 30-2. Work queue 30-1 includes multiple work request entries 60, including work request entry 60-2 and work request entry 60-1 without timing data. Work request entry 60-2 includes timing data 62. Work request entry 60-2 defines the execution timing for the corresponding work request included in that work request entry 60-2. In other words, each work request entry 60-2 includes timing data 62, which defines the execution of the work request included in that work request entry 60-2. Work request entries 60-1 without timing data are typically executed as soon as they are read from work queue 30-1, without being subject to execution time constraints.
[0064] Work queue 30-2 includes work request entries 60-3 and work request entries 60-4 with timing data. Work request entry 60-3 is a fenced entry with timing data 62 defining the execution timing of all work request entries 60-4 shown in work queue 30-2. Work request entries 60-4 include their respective work requests. Work request entry 60-3 is a fenced entry with timing data 62 but no work request. Work queue 30-2 also includes work request entry 60-5, which is a fenced entry with timing data 60 defining other work request entries 60 in work queue 30-2 (but not in...). Figure 5 The fence entry is the timing data 62 for the execution timing of the fence entry 60-3, 60-5. Therefore, the timing data 62 of the corresponding fence entries 60-3, 60-5 defines the execution timing for the work requests included in other work request entries 60.
[0065] Now for reference Figure 6 , Figure 6 It is used for Figure 1 A schematic diagram of one of the job request entries 60-2 in System 10. Job request entry 60 includes a job request 64 (describing the job to be performed), references 66 to one or more locations in memory 24 where the corresponding workload (i.e., the workload to be processed according to the job request) is stored or should be distributed, and timing data 62 (which includes timing data 68 and operators 69). The job to be performed can include any suitable request, such as a request to process packets sent via network 16, a request to process packets received via network 16, a request to encrypt or decrypt data, a request to calculate a CRC, or a request to calculate a cryptographic hash. Timing data 68 can include one or more times. For example, if operator 69 specifies a range, timing data 68 can include two times defining that range. For example, if operator 69 specifies a pattern, timing data 68 can include two times defining that pattern, such as the first 100 microseconds per second. In many cases, timing data 68 includes a single time. Operator 69 can include any suitable operator, such as conditional operators, comparison operators, less than, greater than, equal to, including patterns (e.g., the first 100 microseconds per second), including periodic patterns, within a range (e.g., between X and Y seconds), and masks, which will be described in more detail below.
[0066] Now for reference Figure 7 , Figure 7 It is used for Figure 1 A schematic diagram of job request fence entry 60-3 in system 10. Job request entry 60-3 includes timing data 68 and operators 69, but does not include job request 64 or reference 66. In some embodiments, job request entry 60-3 may include a zero-value job request, that is, a job request for which there is no job to be performed.
[0067] Now for reference Figure 8 , Figure 8 It includes Figure 2 The flowchart 70 shows the steps in the operation method of the peripheral device 14.
[0068] Processing circuitry 32 is configured to read a given work request entry 60 stored in a plurality of work request entries 60 in work queue 30 of memory 24 (block 72). A given work request entry 60 may include timing data 68 and an operator 69. Timing data 68 indicates the time at which at least one work request should be executed. Processing circuitry 32 is configured to retrieve a clock value from hardware clock 18 (block 74). In some embodiments, processing circuitry 32 may include a scheduler (not shown) configured to copy work queue 30 from memory 24 and pass control to an execution engine (not shown) of processing circuitry 32, which reads the given work request entry 60 and executes the following reference. Figure 8 The steps described.
[0069] In decision block 76, processing circuitry 32 is configured to compare timing data 68 included in a given job request entry 60 with a retrieved clock value based on the operator 69 included in the given job request entry 60. If the condition specified by the timing data 68 and operator 69 does not match the retrieved clock value (branch 78), the given job request is returned to queue 30 (or returned to the scheduler), and the next job request entry 60 is read from queue 30 (block 80), and steps 74 and 76 are repeated with the next job request entry 60. For example, if operator 69 is "greater than" and timing data 68 equals 3, but the currently retrieved clock value is 2, then the condition does not match. If the condition does match (branch 82), processing continues to the steps in block 84. For example, if operator 69 is "greater than" and timing data 68 equals 3, and the currently retrieved clock value is 4, then the condition does match. In some embodiments, job request entries 60 may be ordered by execution time. See also Figure 9 A more detailed description of the processing of operator 69, including the mask.
[0070] Processing circuitry 32 is configured to execute a work request associated with a given work request entry 60 using the workload associated with the work request, while timing the execution of the work request in response to timing data 68 and operator 69 (included in the given work request entry) and a retrieved clock value (block 84). In some embodiments, work request 64 and reference 66 are included in the given work request entry 60. In some embodiments, the given work request entry 60 is a fenced entry that does not include work request 64 and reference 66, but provides timing data 62 for other work request entries 60 in the work queue 30 as described above.
[0071] The execution of a job request can include any suitable processing request. Refer to the steps in boxes 86-96 below for some examples.
[0072] Processing circuitry 32 can be configured to process packets (box 86) to be transmitted via network 16 in response to corresponding timing data 68 and corresponding operators 69 of the corresponding entry in work request entry 60. Network interface 36 is configured to transmit packets via network 16.
[0073] Processing circuit 32 can be configured to process corresponding timing data 68 in response to the corresponding work request entry 60 and the corresponding clock value retrieved from hardware clock 18, according to time-division multiplexing schedule 20. Figure 1 Packets that need to be sent via network 16 at regular intervals.
[0074] In some embodiments, network interface 36 is configured to receive packets via network 16, and processing circuitry 32 is configured to process received packets (block 88) timed in response to corresponding timing data 68 and corresponding operator 69 of a corresponding job request entry 60. For example, when a packet is received, it enters a packet processing pipeline of processing circuitry 32. The pipeline may include a guiding decision process. The result of the guiding decision process may include which job request entry the packet belongs to. For example, if the result of the guiding decision process is job request entry number 5, then processing circuitry 32 reads job request entry number 5. If the timing data 68 and operator 69 of the read job request indicate that the packet will not be processed until 9:00 and the current time is 8:59, then the packet may be discarded.
[0075] In some embodiments, the hardware accelerator 40 is configured to perform any one or more of the following: encrypting the workload (box 90); decrypting the workload (box 92); calculating the cyclic redundancy check (CRC) of the workload (box 94); or calculating the cryptographic hash of the workload (box 96).
[0076] Now for reference Figure 9 It is used for Figure 1A schematic diagram of clock value masking in System 10. Box 98 shows the UTC time format, including year (yyyy), month (MM), day (dd), hour (HH), minute (MM), second (ss), and millisecond (SSS). For example, if a workload (e.g., a group) in one of the work queues 30 should be processed within the first 100 milliseconds per second, a mask can be used as operator 69 to specify the bits of the retrieved clock value that should be masked.
[0077] Figure 9 Three distinct time values 100, 104, and 106 are shown, where the most significant bits (year, month, day, hour, minute, and second) of clock values 100, 104, and 106 are masked using mask 102, leaving only the millisecond values (SSS) of clock values 100, 104, and 106. Therefore, the masked millisecond values (SSS) of clock values 100, 104, and 106 can be compared with timing data 68 to determine whether clock values 100, 104, and 106 indicate the first 100 milliseconds of a second, thus revealing clock values 100 and 106 that meet this condition. In some embodiments, the mask may also mask some least significant bits (SS), leaving only one of the digits of the millisecond (SSS) value (the most significant bit, S), so if this digit equals 0, it is known that the clock value is within the first 100 milliseconds of a second.
[0078] The above can be used in a time-division multiplexing system, where packets in queue A are transmitted within the first 100 milliseconds of each second, and packets in queue B are transmitted within the last 100 milliseconds of each second, and so on. Queue A may include a fence entry that specifies masking the most significant bit and checks the remaining bits to determine if the value is between 0 and 99. Queue B may include a fence entry that specifies masking the most significant bit and checks the remaining bits to determine if the value is between 100 and 199, and so on.
[0079] For clarity, the various features of the invention described in the context of a single embodiment may also be provided in combination in a single embodiment. Conversely, for brevity, the various features of the invention described in the context of a single embodiment may also be provided individually or in any suitable sub-combination.
[0080] The embodiments described above are cited by way of example, and the invention is not limited to what has been specifically shown and described above. Rather, the scope of the invention includes combinations and sub-combinations of the various features described above, as well as variations and modifications thereto that would occur to those skilled in the art upon reading the foregoing description and which are not disclosed in the prior art.
Claims
1. A system including peripheral devices, the peripheral devices comprising: Hardware clock; as well as Processing circuitry, used for: Read a given job request entry stored in a plurality of job request entries in at least one work queue in memory, the given job request entry including timing data and an operator, the timing data indicating the time when the job request should be executed, wherein the operator is a conditional operator or a comparison operator; Retrieve the clock value from the hardware clock; as well as The work request with the workload is executed when the execution of the work request is timed in response to the timing data, the operator, and the retrieved clock value. The processing circuitry includes a grouping pipeline for performing a guided decision-making process based on groups associated with the job request entries.
2. The system according to claim 1 further includes a plurality of host devices, each host device comprising: An interface for connecting to and transmitting data with the peripheral device; as well as A processor for running at least one software application configured to write host-specific job request entries from the job request entries into at least one corresponding job queue in the at least one job queue.
3. The system according to claim 1 further includes a host device, the host device comprising: An interface for connecting to and transmitting data with the peripheral device; as well as A processor for running at least one software application configured to write the job request entries into the at least one job queue in the memory, wherein the peripheral device includes an interface configured to connect to and transmit data with the host device.
4. The system of claim 3, wherein the host device includes the memory.
5. The system of claim 3, wherein the at least one software application is configured to write the timing data in Coordinated Universal Time (UTC) format.
6. The system of claim 3, wherein the at least one software application is configured to generate corresponding timing data for writing to the corresponding entry in the work request entry of the at least one work queue in the memory in response to time-division multiplexing scheduling.
7. The system of claim 6, wherein the processing circuitry is configured to process corresponding timing data in response to a corresponding work request entry in the work request entry and a corresponding clock value retrieved from the hardware clock, and packets to be transmitted over the network, timed according to the time-division multiplexing schedule.
8. The system of claim 1, wherein the operator is selected from any one or more of the following: less than; greater than; equal to; including pattern; including periodic pattern; within a range; conforming mask.
9. The system of claim 8, wherein the mask includes masking the most significant bit of the retrieved clock value.
10. The system of claim 1, wherein the corresponding timing data of the corresponding work request entry in the work request entry defines the execution timing of the work request included in other work request entries of the work request entry.
11. The system of claim 10, wherein the corresponding entry in the work request entry is a fence entry with no work request.
12. The system of claim 1, wherein the corresponding work request entry in the work request entry defines the execution timing of the corresponding work request included in the corresponding work request entry in the work request entry.
13. The system of claim 1, wherein at least some of the job request entries include references to locations in the memory where the corresponding workload should be stored or distributed.
14. The system of claim 1, wherein the peripheral device includes a network interface controller and a network interface for sharing packets with remote devices via a packet data network.
15. The system of claim 14, wherein the peripheral device includes a data processing unit (DPU), the data processing unit including the network interface controller, the network interface, and at least a processing core.
16. The system according to claim 14, wherein: The network interface is configured to receive packets through the network; and The processing circuit is configured to process received packets in response to the timing data and corresponding operators of the corresponding work request entries in the work request entries.
17. The system according to claim 14, wherein: The processing circuitry is configured to process packets to be transmitted over the network in response to corresponding timing data and corresponding operators of the corresponding work request entries in the work request entries; and The network interface is configured to send the packets through the network.
18. The system of claim 1, wherein the processing circuitry includes a clock controller for training the hardware clock in response to a Precision Time Protocol (PTP).
19. The system of claim 1, wherein the processing circuitry includes a hardware accelerator for executing at least one of the work requests.
20. The system of claim 19, wherein the hardware accelerator is configured to perform any one or more of the following: Encrypt the workload; Decrypt the workload; Calculate the Cyclic Redundancy Check (CRC) of the workload; or Calculate the cryptographic hash of the workload.
21. A system comprising: Multiple peripheral devices for interconnecting with each other via a network, and including corresponding hardware clocks to be synchronized with each other via the network; as well as Multiple host devices connected to the network via the peripheral devices, each host device including a corresponding processor for running corresponding software applications to share data with each other over the network, each of the corresponding software applications being configured to write a corresponding job request entry to a corresponding job queue, the corresponding job request entry including a corresponding operator and corresponding timing data, the timing data indicating the time when the corresponding job request should be executed, wherein the operator is a conditional operator or a comparison operator, and each of the peripheral devices including processing circuitry for: Read the corresponding job request entries from the corresponding job queue; Retrieve the corresponding clock value from the corresponding hardware clock in the hardware clock; as well as Execute the corresponding work request with workload in response to the corresponding timing data, the corresponding operator, and the retrieved corresponding clock value. The processing circuitry includes a grouping pipeline for performing a guided decision-making process based on groups associated with the job request entries.
22. The system of claim 21, wherein the respective software applications are configured to share the data with each other over the network and generate corresponding job request entries in response to time-division multiplexing scheduling.
23. A method comprising: Read a given job request entry stored in a plurality of job request entries in at least one work queue in memory, the given job request entry including timing data and an operator, the timing data indicating the time when the job request should be executed, wherein the operator is a conditional operator or a comparison operator; Retrieve clock value from hardware clock; as well as The work request is executed when the execution of the work request is timed in response to the timing data, the operator, and the retrieved clock value, wherein a guided decision process is performed using a grouping pipeline based on groups associated with the work request entry.
Citation Information
Patent Citations
Techniques to operate a time division mulitplexing(TDM) media access control (MAC)
US20190319730A1
TDMA Networking using Commodity NIC / Switch
US20210297151A1
Streaming System
US20190379714A1