A method for real-time optimization of data stream computing bandwidth
By using ring buffers and registers to record the number and time of instruction times and adjusting the bandwidth allocation strategy in real time, the problem of unreasonable hardware bandwidth caused by the ratio differences between fixed-point and floating-point instructions is solved, and the computing efficiency and bandwidth utilization are improved.
Patent Information
- Application Number
- CN202411592559.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-08
AI Technical Summary
In traditional data flow calculation, the hardware bandwidth allocation caused by the difference in proportion to fixed-point instructions and floating-point instructions is unreasonable, resulting in low computing efficiency and insufficient bandwidth.
The number and time of write and output of operation instructions are recorded through the ring buffer and register, and the bandwidth allocation strategy is adjusted in real time with the bandwidth judge to ensure the reasonable allocation of bandwidth requirements for fixed-point and floating-point operations.
It improves memory utilization efficiency and security, alleviates the problem of insufficient bandwidth, improves the fluency and real-timeness of data processing, and enhances the adaptability and robustness of the system.
Smart Images

Figure CN119520429B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high-performance computing in computers. Specifically, it relates to a method for real-time optimization of data stream computing bandwidth. Background Art
[0002] High-performance computing aims to process multiple tasks, instructions, or data items simultaneously. Its implementation relies on a high-performance computer system that organizes thousands of processors in a specific order through network connections. This organization involves not only the interconnection topology and communication protocols of the network but also levels such as the operating system and middleware software. The core goal of high-performance computing is to provide a computing speed beyond that of traditional computers and solve complex problems that traditional computers cannot handle.
[0003] In the computational models of computer architectures, control flow and data flow are the two main classifications. Control-flow computers, i.e., von Neumann-type computers, as the mainstream architecture, follow instruction-sequential-driven operations, and the participation of data operations depends on the requirements of the current instruction. In contrast, the data flow computational model adopts a data-driven approach, where the execution of an instruction depends on the completeness of its required operands, and the execution result further drives the execution of subsequent instructions.
[0004] A data flow computing system usually consists of multiple computing nodes. Each node has powerful computing capabilities but relatively weak control capabilities and low complexity. These nodes contain multiple fixed-point arithmetic units and floating-point arithmetic units, corresponding to fixed-point arithmetic pipelines and floating-point arithmetic pipelines respectively. In each clock cycle, the pipeline selects an instruction from the ready instructions for execution. However, the ratio of fixed-point instructions to floating-point instructions varies in different computing applications, which directly affects the requirements for fixed-point arithmetic and floating-point arithmetic bandwidths.
[0005] Inside the nodes of data flow computing, the instruction emission mechanism plays a key role. As Figure 3 shown, the instruction queue is used to store the real-time execution information of instructions, and the instruction emission logic is responsible for selecting and emitting instructions to the corresponding arithmetic units. However, the traditional instruction emission method has certain limitations: when the proportion of fixed-point instructions is high and the number of floating-point instructions is small, the computing pipeline of the floating-point arithmetic unit may be idle, while the bandwidth of the fixed-point pipeline may be tight. This fixed and non-adjustable hardware bandwidth allocation method not only reduces the execution efficiency of computing but may also lead to insufficient bandwidth when the data volume is too large, thereby affecting the normal transmission of data and the timely response of computing requests, and ultimately reducing the computing processing speed.
[0006] Regarding the problems in the related art, no effective solutions have been proposed yet. Summary of the Invention
[0007] In view of the problems in the related art, the present invention proposes a method for real-time optimization of the data stream calculation bandwidth to overcome the above-mentioned technical problems existing in the existing related art.
[0008] To this end, the specific technical solution adopted by the present invention is as follows:
[0009] A method for real-time optimization of the data stream calculation bandwidth, the method comprising:
[0010] S1. Extract arithmetic instructions from the instruction queue of the data stream, and record the write times and write times of each type of arithmetic instruction through a counter in the circular buffer, and transmit the write times and write times to the bandwidth discriminator for the first bandwidth judgment;
[0011] S2. Transmit the arithmetic instructions in the circular buffer to the register, and record the output times and output times of each type of arithmetic instruction through a counter in the register, and transmit the output times and output times to the bandwidth discriminator for the second bandwidth judgment;
[0012] S3. The processor component adjusts the bandwidth allocation strategy in real time according to the results of the first and second bandwidth judgments, combined with the proportion of the arithmetic instruction types. The processor component includes a data processor and a bandwidth allocator;
[0013] S4. Based on the adjusted bandwidth allocation strategy, transmit the arithmetic instructions to the corresponding arithmetic unit to execute the corresponding arithmetic instructions.
[0014] Further, extracting arithmetic instructions from the instruction queue of the data stream, and recording the write times and write times of each type of arithmetic instruction through a counter in the circular buffer, and transmitting the write times and write times to the bandwidth discriminator for the first bandwidth judgment includes:
[0015] S11. Obtain the instruction queue of the data stream, and allocate the instruction queue to the arithmetic pipeline. The arithmetic pipeline includes a fixed-point arithmetic pipeline and a floating-point arithmetic pipeline;
[0016] S12. Extract fixed-point arithmetic instructions and floating-point arithmetic instructions from the fixed-point arithmetic pipeline and the floating-point arithmetic pipeline respectively;
[0017] S13. Store the extracted fixed-point arithmetic instructions and floating-point arithmetic instructions into the corresponding circular buffers respectively, and use the counters in the circular buffers to record the write times and write times. The write times include the fixed-point write times and the floating-point write times, and the write times include the fixed-point write time and the floating-point write time;
[0018] S14. Transmit the recorded write times and write times to the bandwidth discriminator, and use the bandwidth discriminator to perform the first bandwidth judgment.
[0019] Further, obtain the instruction queue of the data stream and allocate the instruction queue to the operation pipeline. The operation pipeline includes a fixed-point operation pipeline and a floating-point operation pipeline, and it includes:
[0020] S111. Obtain the instruction queue of the data stream and check the status of the instructions in the instruction queue;
[0021] S112. Based on the instructions after the check, obtain and judge the status of the fixed-point operation pipeline and the floating-point operation pipeline;
[0022] S113. Based on the judgment results of the fixed-point operation pipeline and the floating-point operation pipeline, transfer the instructions in the instruction queue to the instruction issue logic;
[0023] S114. The instruction issue logic selects the corresponding pipeline according to the type of the instruction, removes the corresponding instruction from the instruction queue, and allocates it to the selected pipeline;
[0024] S115. The fixed-point operation pipeline and the floating-point operation pipeline respectively receive and execute the allocated instructions, and update the status of the instruction issue logic based on the execution results.
[0025] Further, transfer the recorded write times and write times to the bandwidth judge, and the initial bandwidth judgment by using the bandwidth judge includes:
[0026] S141. Transfer the recorded fixed-point write times, floating-point write times, fixed-point write times, and floating-point write times to the bandwidth judge;
[0027] S142. Based on the transfer result, the bandwidth judge calculates the total write bandwidth by using the bandwidth formula;
[0028] S143. Compare the calculated total write bandwidth with the node bandwidth to perform the initial bandwidth judgment.
[0029] Further, the bandwidth formula is:
[0030]
[0031] In the formula, Z represents the total write bandwidth value;
[0032] w1 represents the fixed-point write times;
[0033] t1 represents the fixed-point write time;
[0034] w2 represents the floating-point write times;
[0035] t2 represents the floating-point write time.
[0036] Further, comparing the calculated total write bandwidth with the node bandwidth to perform the initial bandwidth judgment includes:
[0037] If the total write bandwidth is less than or equal to the node bandwidth, the bandwidth judge does not send an adjustment instruction to the clock regulator;
[0038] If the total write bandwidth is greater than the node bandwidth, the bandwidth judge sends an instruction to the clock regulator and increases the clock cycle according to a preset period.
[0039] Furthermore, transferring the arithmetic instructions in the circular buffer to the register, recording the output times and output times of each type of arithmetic instruction through the counter in the register, and transmitting the output times and output times to the bandwidth judge for secondary bandwidth judgment includes:
[0040] S21. Read the fixed-point arithmetic instruction and floating-point arithmetic instruction from the circular buffer and transfer the fixed-point arithmetic instruction and floating-point arithmetic instruction to the corresponding register;
[0041] S22. Use the counter in the register to record the output times and output times. The output times include the fixed-point output times and floating-point output times, and the output times include the fixed-point output time and floating-point output time;
[0042] S23. Feed back the recorded output times and output times to the bandwidth judge and use the bandwidth judge to perform secondary bandwidth judgment.
[0043] Furthermore, feeding back the recorded output times and output times to the bandwidth judge and using the bandwidth judge to perform secondary bandwidth judgment includes:
[0044] S231. Feed back the recorded fixed-point output times, floating-point output times, fixed-point output time, and floating-point output time to the bandwidth judge;
[0045] S232. Based on the feedback result, the bandwidth judge calculates the total output bandwidth using the bandwidth formula;
[0046] S233. Compare the calculated total output bandwidth with the node bandwidth to perform secondary bandwidth judgment.
[0047] Furthermore, comparing the calculated total output bandwidth with the node bandwidth to perform secondary bandwidth judgment includes:
[0048] If the total output bandwidth is less than or equal to the node bandwidth, the bandwidth judge does not send an adjustment instruction to the clock regulator;
[0049] If the total output bandwidth is greater than the node bandwidth, the bandwidth judge sends an instruction to the clock regulator and increases the clock cycle according to a preset period.
[0050] Further, the processor component adjusts the bandwidth allocation strategy in real time according to the results of the primary and secondary bandwidth judgments, combined with the ratio of the types of arithmetic instructions. The processor component includes a data processor and a bandwidth allocator, including:
[0051] S31. Transmit the primary bandwidth judgment result and the secondary bandwidth judgment result from the register to the data processor;
[0052] S32. The data processor receives and processes the fixed-point arithmetic instructions and floating-point arithmetic instructions in the primary bandwidth judgment result and the secondary bandwidth judgment result, obtains the fixed-point processing result and the floating-point processing result, and transmits the fixed-point processing result and the floating-point processing result to the bandwidth allocator;
[0053] S33. The bandwidth allocator adjusts the bandwidth allocation strategy according to the ratio of the received fixed-point processing result and floating-point processing result.
[0054] The beneficial effects of the present invention are as follows:
[0055] 1. By using a circular buffer to cache data, the present invention avoids frequent memory allocation operations, reduces the overhead of memory management, improves the utilization efficiency and security of memory; at the same time, the circular buffer plays a role of data buffering and temporary storage, effectively alleviates the problem of insufficient bandwidth caused by excessive data volume, avoids data backlog and slow processing, and improves the fluency and real-time performance of data processing.
[0056] 2. The present invention adjusts the bandwidth ratio occupied by fixed-point arithmetic and floating-point arithmetic in real time according to the ratio of the number of register outputs, makes full use of bandwidth resources, and improves bandwidth utilization; in addition, by judging the number of write operations of the circular buffer, the total bandwidth required for the current instruction is determined, ensuring that the bandwidth can be reasonably allocated when processing a large amount of data, and avoiding problems such as data overflow and low arithmetic efficiency caused by insufficient bandwidth.
[0057] 3. By optimizing bandwidth allocation from two dimensions of time and space, the present invention effectively solves the problem of low computing efficiency caused by tight bandwidth allocation in the data flow architecture. By dynamically adjusting the bandwidth allocation strategy, the efficient operation of the system is ensured, and the waste of bandwidth resources is avoided; in addition, the real-time optimization method can flexibly respond to different computing tasks and data traffic, improving the adaptability and robustness of the system. Description of the Drawings
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0059] Figure 1 is a flowchart of a method for real-time optimization of data stream computing bandwidth according to an embodiment of the present invention;
[0060] Figure 2 is a schematic diagram of the instruction emission structure of the data stream architecture in a method for real-time optimization of data stream computing bandwidth according to an embodiment of the present invention;
[0061] Figure 3 is a schematic diagram of the instruction emission structure of the traditional data stream architecture in a method for real-time optimization of data stream computing bandwidth according to an embodiment of the present invention. Detailed implementation manners
[0062] To further illustrate the embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.
[0063] According to an embodiment of the present invention, a method for real-time optimization of data stream computing bandwidth is provided.
[0064] Now, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners. As Figure 1 shown, the method for real-time optimization of data stream computing bandwidth according to an embodiment of the present invention includes:
[0065] S1. Extract arithmetic instructions from the instruction queue of the data stream, and record the write times and write counts of each type of arithmetic instruction through a counter in the circular buffer, and transmit the write counts and write times to the bandwidth judge for primary bandwidth judgment.
[0066] Specifically, extracting arithmetic instructions from the instruction queue of the data stream, and recording the write times and write counts of each type of arithmetic instruction through a counter in the circular buffer, and transmitting the write counts and write times to the bandwidth judge for primary bandwidth judgment includes:
[0067] S11. Obtain the instruction queue of the data stream, and allocate the instruction queue to the arithmetic pipeline, where the arithmetic pipeline includes a fixed-point arithmetic pipeline and a floating-point arithmetic pipeline.
[0068] Specifically, obtaining the instruction queue of the data stream, and allocating the instruction queue to the arithmetic pipeline, where the arithmetic pipeline includes a fixed-point arithmetic pipeline and a floating-point arithmetic pipeline includes:
[0069] S111. Obtain the instruction queue of the data stream, and check the status of the instructions in the instruction queue;
[0070] S112. Based on the instructions after the inspection is completed, obtain and judge the status of the fixed-point operation pipeline and the floating-point operation pipeline;
[0071] S113. Based on the judgment results of the fixed-point operation pipeline and the floating-point operation pipeline, transfer the instructions in the instruction queue to the instruction issue logic;
[0072] S114. The instruction issue logic selects the corresponding pipeline according to the type of the instruction, removes the corresponding instruction from the instruction queue, and allocates it to the selected pipeline;
[0073] S115. The fixed-point operation pipeline and the floating-point operation pipeline respectively receive and execute the allocated instructions, and update the status of the instruction issue logic based on the execution results.
[0074] S12. Extract the fixed-point operation instructions and the floating-point operation instructions from the fixed-point operation pipeline and the floating-point operation pipeline respectively.
[0075] S13. Store the extracted fixed-point operation instructions and floating-point operation instructions into the corresponding circular buffers respectively, and use the counters in the circular buffers to record the write times and write times, where the write times include the fixed-point write times and the floating-point write times, and the write times include the fixed-point write times and the floating-point write times.
[0076] S14. Transmit the recorded write times and write times to the bandwidth judge, and use the bandwidth judge to perform the initial bandwidth judgment.
[0077] Specifically, transmitting the recorded write times and write times to the bandwidth judge and using the bandwidth judge to perform the initial bandwidth judgment includes:
[0078] S141. Transmit the recorded fixed-point write times, floating-point write times, fixed-point write times, and floating-point write times to the bandwidth judge.
[0079] S142. Based on the transmission results, the bandwidth judge calculates the total write bandwidth using the bandwidth formula.
[0080] Specifically, the bandwidth formula is:
[0081]
[0082] In the formula, Z represents the total write bandwidth value;
[0083] w1 represents the fixed-point write times;
[0084] t1 represents the fixed-point write time;
[0085] w2 represents the floating-point write times;
[0086] t2 represents the floating-point write time.
[0087] S143. Compare the calculated total write bandwidth with the node bandwidth to perform an initial bandwidth judgment.
[0088] Specifically, comparing the calculated total write bandwidth with the node bandwidth to perform an initial bandwidth judgment includes:
[0089] If the total write bandwidth is less than or equal to the node bandwidth, the bandwidth judge does not send an adjustment instruction to the clock regulator.
[0090] If the total write bandwidth is greater than the node bandwidth, the bandwidth judge sends an instruction to the clock regulator and increases the clock cycle according to a preset period.
[0091] It should be added that the number of fixed-point write times and floating-point write times in the counter result are passed to the bandwidth judge, and it is judged through calculation whether the total bandwidth of fixed-point operation and floating-point operation is greater than the actual node bandwidth. If so, the bandwidth judge issues an instruction and the clock regulator responds, and the clock cycle continuously increases in units of T cycle. If not, the bandwidth judge does not issue an instruction and the clock regulator does not respond.
[0092] S2. Transmit the operation instructions in the circular buffer to the register, record the output times and output times of each type of operation instruction through the counter in the register, and transmit the output times and output times to the bandwidth judge for a secondary bandwidth judgment.
[0093] Specifically, transmitting the operation instructions in the circular buffer to the register, recording the output times and output times of each type of operation instruction through the counter in the register, and transmitting the output times and output times to the bandwidth judge for a secondary bandwidth judgment includes:
[0094] S21. Read the fixed-point operation instructions and floating-point operation instructions from the circular buffer and transmit the fixed-point operation instructions and floating-point operation instructions to the corresponding registers.
[0095] It should be added that the speed at which the fixed-point operation instructions and floating-point operation instructions are cached in the circular buffer is greater than the speed at which their data is obtained from the circular buffer.
[0096] S22. Use the counter in the register to record the output times and output times. The output times include fixed-point output times and floating-point output times, and the output times include fixed-point output times and floating-point output times.
[0097] S23. Feed back the recorded output times and output times to the bandwidth judge and use the bandwidth judge to perform a secondary bandwidth judgment.
[0098] Specifically, the recorded output times and output time are fed back to the bandwidth judge, and the secondary bandwidth judgment by the bandwidth judge includes:
[0099] S231. Feed back the recorded fixed-point output times, floating-point output times, fixed-point output time, and floating-point output time to the bandwidth judge;
[0100] S232. Based on the feedback result, the bandwidth judge calculates the total output bandwidth using the bandwidth formula;
[0101] S233. Compare the calculated total output bandwidth with the node bandwidth for secondary bandwidth judgment.
[0102] Specifically, comparing the calculated total output bandwidth with the node bandwidth for secondary bandwidth judgment includes:
[0103] If the total output bandwidth is less than or equal to the node bandwidth, the bandwidth judge does not send an adjustment instruction to the clock regulator;
[0104] If the total output bandwidth is greater than the node bandwidth, the bandwidth judge sends an instruction to the clock regulator and increases the clock cycle according to a preset period.
[0105] It should be added that the initial clock cycle of the register is very small, and the delay time of data in the register can be ignored. If the clock regulator runs, the register clock cycle increases continuously in T cycles, and its delay will also increase gradually.
[0106] S3. The processor component adjusts the bandwidth allocation strategy in real time according to the results of the primary and secondary bandwidth judgments and in combination with the proportion of the operation instruction types. The processor component includes a data processor and a bandwidth allocator.
[0107] Specifically, the processor component adjusts the bandwidth allocation strategy in real time according to the results of the primary and secondary bandwidth judgments and in combination with the proportion of the operation instruction types. The processor component includes a data processor and a bandwidth allocator, including:
[0108] S31. Transmit the primary bandwidth judgment result and the secondary bandwidth judgment result from the register to the data processor.
[0109] It should be added that the fixed-point operation instructions, floating-point operation instructions, fixed-point output times, and floating-point output times in the register are transmitted to the data processor.
[0110] S32. The data processor receives and processes the fixed-point operation instructions and floating-point operation instructions in the primary bandwidth judgment result and the secondary bandwidth judgment result to obtain the fixed-point processing result and the floating-point processing result, and transmits the fixed-point processing result and the floating-point processing result to the bandwidth allocator.
[0111] S33. The bandwidth allocator adjusts the bandwidth allocation strategy according to the ratio of the fixed-point processing result and the floating-point processing result received.
[0112] S4. Based on the adjusted bandwidth allocation strategy, the arithmetic instructions are transmitted to the corresponding arithmetic units to execute the corresponding arithmetic instructions.
[0113] It should be added that after each group of instruction queues is processed, each counter is cleared, and the clock regulator is also reset, waiting for the next group of instruction queues.
[0114] In a specific embodiment, the instruction queue is used to store the real-time execution information of instructions, and a group of instruction queues waits for arithmetic operations. The queue enters the instruction issue logic, and the instruction issue logic is used to select instructions and issue them into the corresponding arithmetic pipelines. In each clock cycle, the instruction issue logic selects an instruction from the ready instructions and issues it to the corresponding pipeline for execution. The ratio of fixed-point arithmetic instructions and floating-point arithmetic instructions in different computing applications is also different.
[0115] As Figure 2 shown, 2 arithmetic pipelines are set up, namely the fixed-point arithmetic pipeline and the floating-point arithmetic pipeline. The instruction issue logic first judges whether the status of each instruction in the instruction queue is ready, and whether each arithmetic pipeline can receive and execute new instructions. If the judgment is yes, the instructions in the instruction queue are allocated to the corresponding pipelines for execution.
[0116] The instructions of the fixed-point arithmetic pipeline enter the circular buffer 1, and the instructions of the floating-point arithmetic pipeline enter the circular buffer 2; counters counter1 and counter2 are respectively set at the data writing places of the circular buffer 1 and the circular buffer 2. All fixed-point arithmetic instructions and floating-point arithmetic instructions are written into the circular buffer, and their counting results are w1 and w2 respectively, and the time used is t1 and t2 respectively; the counters send the counting results w1 and w2 to the bandwidth judge, and the bandwidth judge makes a judgment, where the arithmetic bandwidth of this node is B.
[0117] If w1 / t1 + w2 / t2 <= B, the bandwidth judge judges it as no, does not send instructions, and the clock regulator does not operate; the following descriptions all occur under this condition.
[0118] The fixed-point arithmetic instructions and the floating-point arithmetic instructions are respectively read out from the circular buffer 1 and the circular buffer 2 and enter the register 1 and the register 2; among them, the circular buffer 1 and the circular buffer 2, and the register 1 and the register 2 are all of the same type. The speed of caching the data in the circular buffer is greater than the speed of obtaining the data from the circular buffer, avoiding data congestion in the circular buffer and improving the data transmission efficiency; the initial clock cycle T0 of the register is very small, and the delay of the data in the register can be ignored.
[0119] Set counters counter3 and counter4 at the output positions of register 1 and register 2. All fixed-point operation instructions and floating-point operation instructions enter the data processor from the registers, and their counting results are O1 and O2 respectively, and the time taken is t3 and t4 respectively. The data of O1 and O2 are sent to the data processor, and the processed data results are O1 / t3 and O2 / t4. The initial values of the bandwidths of the fixed-point operation unit and the floating-point operation unit are both B / 2. Send the processed results to the bandwidth allocator, and the allocator pre-allocates the bandwidth in the ratio of O1 / t3:O2 / t4. The fixed-point operation instructions and floating-point operation instructions enter the fixed-point operation unit and the floating-point operation unit respectively for data operation processing. When the next instruction queue arrives, all counters are set to 0, the register clock cycle is reset to the initial value, and the bandwidths of the fixed-point operation unit and the floating-point operation unit are also reset.
[0120] If w1 / t1 + w2 / t2 > B, the bandwidth judge determines it as yes, sends an instruction, the clock regulator responds, the register cycle is adjusted to T1, and O1 / (t3 + T1) + O2 / (t3 + T1) <= B, from which T1 can be obtained. The following descriptions all occur under this condition.
[0121] The fixed-point operation instructions and floating-point operation instructions are read out from circular buffer 1 and circular buffer 2 respectively and enter register 1 and register 2. At this time, the clock regulator is already in the response state, the register cycle is adjusted to T1, and counter3 and counter4 output O1 and O2. At the same time, the data of O1 and O2 are sent to the data processor, and the processed data results are O1 / t3 and O2 / t4. Send the processed results to the bandwidth allocator, and the allocator pre-allocates the bandwidth in the ratio of O1 / t3:O2 / t4. The fixed-point operation instructions and floating-point operation instructions enter the fixed-point operation unit and the floating-point operation unit respectively for data operation processing.
[0122] In addition, for fixed-point numbers, the decimal point is not fixed. For example, for int, the decimal point is at the last digit of the number. The more decimal digits, the higher the precision of the number. If there are n digits after the decimal point, the precision is 1 / (2 n ). The more integer digits, the larger the maximum value that can be represented. The calculation formula for fixed-point numbers is:
[0123]
[0124] In the formula, n represents a decimal number; fc represents the length of the decimal part; bw represents the width of the fixed-point number; i represents an index; Bi represents the value of the i-th bit.
[0125] Floating-point numbers are numbers with an unfixed decimal point, such as float and double. Their representation logic is completely different from that of unsigned integers like int and char. Floating-point numbers cannot be shifted because bits are in different fields and have different meanings. For example:
[0126]
[0127] In the formula, V represents a floating-point number; S represents the sign bit, which can take values of 0 or 1 and determines the sign of a number. 0 indicates positive and 1 indicates negative; M represents the mantissa, which is expressed as a decimal. For example, for 1.234*10 0 , 1.234 is the mantissa; R represents the radix. For a decimal number, R is 10, and for a binary number, R is 2; E represents the exponent, which is expressed as an integer. For example, for 10 -1 , -1 is the exponent.
[0128] If you want to represent a number as a floating-point number in a computer, you only need to confirm these variables.
[0129] In summary, by means of the above technical solution of the present invention, by using a circular buffer to cache data, frequent memory allocation operations are avoided, the overhead of memory management is reduced, and the utilization efficiency and security of memory are improved; at the same time, the circular buffer plays a role in buffering and temporarily storing data, effectively alleviating the problem of insufficient bandwidth caused by excessive data volume, avoiding data backlog and slow processing, and improving the fluency and real-time performance of data processing. By adjusting the bandwidth ratio occupied by fixed-point operations and floating-point operations in real time according to the ratio of the number of register outputs, the bandwidth resources are fully utilized and the bandwidth utilization rate is improved; in addition, by judging the total bandwidth required for the current instruction based on the number of write operations of the circular buffer, it is ensured that when processing a large amount of data, the bandwidth can be reasonably allocated, avoiding data overflow and low operation efficiency problems caused by insufficient bandwidth. By optimizing the bandwidth allocation from two dimensions of time and space, the problem of low computing efficiency caused by tight bandwidth allocation in the data flow architecture is effectively solved. By dynamically adjusting the bandwidth allocation strategy, the efficient operation of the system is ensured and the waste of bandwidth resources is avoided; in addition, the real-time optimization method can flexibly handle different computing tasks and data traffic, improving the adaptability and robustness of the system.
[0130] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for real-time optimization of data flow computing bandwidth, characterized in that: The method includes: S1, extracting operation instructions from the instruction queue of the data stream, and recording the number of writes and the write time of each type of operation instruction through the counter in the ring buffer, and transmitting the number of writes and the write time to the bandwidth determiner for initial bandwidth determination; S2, transferring the operation instructions in the ring buffer to the register, and recording the output times and output time of each type of operation instructions through the counter in the register, and transferring the output times and output time to the bandwidth determiner for secondary bandwidth determination; S3, the processor component adjusts the bandwidth allocation strategy in real time according to the results of the initial and secondary bandwidth judgments and the proportion of the operation instruction types, and the processor component includes a data processor and a bandwidth allocator; S4, based on the adjusted bandwidth allocation strategy, transmitting the operation instruction to the corresponding operation unit, and executing the corresponding operation instruction; The S3 includes: S31, transmitting the initial bandwidth determination result and the secondary bandwidth determination result from the register to the data processor; S32, the data processor receives and processes the fixed-point operation instructions and floating-point operation instructions in the initial bandwidth determination result and the secondary bandwidth determination result, obtains the fixed-point processing result and the floating-point processing result, and transmits the fixed-point processing result and the floating-point processing result to the bandwidth allocator; S33. The bandwidth allocator adjusts the bandwidth allocation strategy according to the ratio of the received fixed-point processing result to the received floating-point processing result.
2. A method for real-time optimization of data flow computing bandwidth according to claim 1, characterized in that: The extracting operation instructions from the instruction queue of the data stream, recording the number of writes and the write time of each type of operation instruction through a counter in the ring buffer, and transmitting the number of writes and the write time to the bandwidth determiner for initial bandwidth determination includes: S11, obtaining an instruction queue of a data stream, and assigning the instruction queue to a calculation pipeline, wherein the calculation pipeline includes a fixed-point calculation pipeline and a floating-point calculation pipeline; S12, extracting fixed-point operation instructions and floating-point operation instructions from the fixed-point operation pipeline and the floating-point operation pipeline respectively; S13, storing the extracted fixed-point operation instructions and floating-point operation instructions in corresponding ring buffers respectively, and using a counter in the ring buffer to record the number of writes and the write time, wherein the number of writes includes the number of fixed-point writes and the number of floating-point writes, and the write time includes the fixed-point write time and the floating-point write time; S14: The recorded number of write times and write time are transmitted to a bandwidth determiner, and the bandwidth determiner is used to perform an initial bandwidth determination.
3. A method for real-time optimization of data flow computing bandwidth according to claim 2, characterized in that: The instruction queue of the data stream is obtained, and the instruction queue is allocated to the operation pipeline, wherein the operation pipeline includes a fixed-point operation pipeline and a floating-point operation pipeline, including: S111, obtaining the instruction queue of the data stream, and checking the status of the instructions in the instruction queue; S112, based on the instructions after the inspection is completed, obtaining and determining the states of the fixed-point operation pipeline and the floating-point operation pipeline; S113, based on the determination results of the fixed-point operation pipeline and the floating-point operation pipeline, transmitting the instructions in the instruction queue to the instruction issuing logic; S114, the instruction issuing logic selects a corresponding pipeline according to the type of the instruction, removes the corresponding instruction from the instruction queue, and distributes it to the selected pipeline; S115 , the fixed-point operation pipeline and the floating-point operation pipeline respectively receive and execute the assigned instructions, and update the state of the instruction issuance logic based on the execution results.
4. A method for real-time optimization of data flow computing bandwidth according to claim 3, characterized in that: The step of transmitting the recorded number of write times and write time to the bandwidth determiner and using the bandwidth determiner to perform initial bandwidth determination includes: S141, transmitting the recorded fixed-point writing times, floating-point writing times, fixed-point writing time and floating-point writing time to a bandwidth determiner; S142: Based on the transmission result, the bandwidth determiner calculates the total write bandwidth using a bandwidth formula; S143: Compare the calculated total write bandwidth with the node bandwidth to perform an initial bandwidth determination.
5. A method for real-time optimization of data flow computing bandwidth according to claim 4, characterized in that: The bandwidth formula is: ; In the formula, Z Indicates the total write bandwidth value; w 1 Indicates the number of fixed-point writes; t 1 Indicates the fixed-point writing time; w 2 Indicates the number of floating point writes; t 2 Indicates floating point write time.
6. A method for real-time optimization of data flow computing bandwidth according to claim 5, characterized in that: The step of comparing the calculated total write bandwidth with the node bandwidth to perform initial bandwidth determination includes: If the total written bandwidth is less than or equal to the node bandwidth, the bandwidth determiner does not send an adjustment instruction to the clock adjuster; If the total write bandwidth is greater than the node bandwidth, the bandwidth determiner sends an instruction to the clock adjuster and increases the clock cycle according to a preset cycle.
7. A method for real-time optimization of data flow computing bandwidth according to claim 1, characterized in that: The method of transmitting the operation instructions in the ring buffer to the register, recording the output times and output time of each type of operation instructions by the counter in the register, and transmitting the output times and output time to the bandwidth determiner for secondary bandwidth determination includes: S21, reading fixed-point arithmetic instructions and floating-point arithmetic instructions from the ring buffer, and transferring the fixed-point arithmetic instructions and floating-point arithmetic instructions to corresponding registers; S22, using a counter in a register to record the number of outputs and the output time, the number of outputs includes the number of fixed-point outputs and the number of floating-point outputs, and the output time includes the fixed-point output time and the floating-point output time; S23, feeding back the recorded output times and output time to the bandwidth determiner, and using the bandwidth determiner to perform secondary bandwidth determination.
8. A method for real-time optimization of data flow computing bandwidth according to claim 7, characterized in that: Feeding back the recorded output times and output time to the bandwidth determiner, and using the bandwidth determiner to perform secondary bandwidth determination includes: S231, feeding back the recorded fixed-point output times, floating-point output times, fixed-point output time and floating-point output time to the bandwidth determiner; S232, based on the feedback result, the bandwidth determiner calculates the output total bandwidth using the bandwidth formula; S233: Compare the calculated total output bandwidth with the node bandwidth to perform secondary bandwidth judgment.
9. A method for real-time optimization of data flow computing bandwidth according to claim 8, characterized in that: The step of comparing the calculated total output bandwidth with the node bandwidth and performing secondary bandwidth determination includes: If the total output bandwidth is less than or equal to the node bandwidth, the bandwidth determiner does not send an adjustment instruction to the clock adjuster; If the total output bandwidth is greater than the node bandwidth, the bandwidth determiner sends an instruction to the clock adjuster and increases the clock cycle according to a preset cycle.
Citation Information
Patent Citations
Calculating system and method for dynamically adjusting resource bandwidth of data flow architecture
CN107688471A
Processing method and device for realizing dynamic adaptation of terminal based on annular buffer mechanism in receiver system, processor and storage medium thereof
CN118041413A