A Custom Instruction Design Method for a Coprocessor in the Replacement Process of a Data Fusion Algorithm Based on an Auction Algorithm

By designing coprocessors and custom instructions, the parallelism problem of single-core processors in the replacement process in the auction algorithm is solved, efficient data fusion algorithm calculation is realized, and real-time and efficiency are improved.

CN114721723BActive Publication Date: 2025-07-18SHANGHAI XIAOCHI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210161583.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-07-18
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

When the existing data fusion algorithm based on auction algorithm is operated on a single-core processor, the replacement process needs to read data from the storage through the bus, resulting in the program execution sequentially, and the parallelism cannot be found, which reduces the performance and real-time nature of the algorithm.

Method used

A coprocessor based on the data fusion algorithm of auction algorithm is designed, and the unique bus peripherals and coprocessor data interaction is achieved using the characteristics of dual-port RAM. The interaction between the coprocessor and the main processor is controlled through custom instructions, the algorithm process is optimized by a finite state machine, and the parallel computing is achieved using dual-port cache.

Benefits of technology

It doubles the computational efficiency of the algorithm, reduces the data handling process of the main processor, realizes direct interaction between the coprocessor and the peripheral, and maximizes the computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114721723B_ABST
    Figure CN114721723B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of computer algorithms, and particularly to a coprocessor for the replacement process in a data fusion algorithm based on an auction algorithm. The coprocessor includes a coprocessor, and the internal circuit of the coprocessor contains a control logic and a computing unit. The control logic unit is used to control the data flow control of the entire computing unit and the data reading and writing with the storage cache. The computing unit is responsible for calculating a pair of data transmitted from a custom instruction, comparing whether the values read from two storage caches are the same, and sending the comparison result to the control logic. For an algorithm with repeated and regular operation steps, implementing it using a finite state machine saves a large amount of time compared to implementing it using a processor. At the same time, through reasonable design, the two data retrievals and comparisons of each group of numbers are synchronized, and optimization is carried out at the algorithm level, doubling the efficiency. By utilizing the characteristics of a dual-port cache, a direct interaction method between the peripheral device and the coprocessor is achieved, rather than accessing the on-chip storage through an ROCC interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer algorithms, and particularly to a coprocessor and a custom instruction design method for the replacement process in a data fusion algorithm based on an auction algorithm. Background Art

[0002] The data fusion algorithm based on the auction algorithm is widely applied to multi-target allocation fields such as radar target tracking, target firepower allocation, and UAV swarm command. This algorithm has characteristics such as a large number of target objects, few calculation links, decentralization, and good implementability.

[0003] The data fusion algorithm based on the auction algorithm can be divided into three parts in terms of process. The input of the first part is two sets of three-dimensional Euclidean coordinates, denoted as sets J and K. The first part is used to calculate the distance between any two points in J and K. The second part is used to screen out valid coordinates. It compares all the distance results with a threshold. When a certain distance value is less than the threshold, the serial numbers corresponding to this distance in J and K are recorded. If a serial number in J and K appears only once and the distance between these two serial numbers is less than the threshold, it means that these two points are in one-to-one match. If a serial number appears multiple times, it means that this serial number is associated with multiple coordinates, and these serial numbers are rearranged. The third step is to perform an auction operation on the rearranged serial numbers waiting for auction in J and K. Among them, the serial numbers in J are bidders, and the serial numbers in K are objects to be auctioned. The price of each bidder in J is the Euclidean distance between J and K. Each item is auctioned once, and the profit of each bidder is the current price minus the current bidding price. The serial number in J with the smallest profit is matched. At the same time, each serial number in J can only match one serial number in K. After each match, the previous match results need to be retrieved. If there are duplicate serial numbers in J or K, the corresponding serial number in the other K or J is retrieved and auctioned again. After the auction, the matching results of the second and third steps are integrated together to obtain the final matching result.

[0004] The current problem is that the existing algorithm uses a single-core processor as the operation platform. When the processor operates in the replacement link of the auction algorithm, each replacement needs to read from the memory through the bus, and each bus read has to go through operations such as fetching values, decoding, execution, peripheral response, and storing to memory. Moreover, the program can only be executed sequentially, and its parallelism cannot be found, which greatly reduces the performance of the algorithm and thus affects its real-time performance.

[0005] In view of the disadvantages of the prior art, the present invention proposes the following solutions:

[0006] 1. Design a coprocessor circuit for the replacement module in this algorithm, and use the characteristics of dual-port RAM to design a unique data interaction method between the bus peripheral and the coprocessor to maximize the calculation efficiency.

[0007] 2. By using the built-in ROCC interface of the open-source processor RocketChip, the behavior of the coprocessor is controlled through custom instructions to achieve the interaction between the coprocessor and the main processor. Summary of the Invention

[0008] (1) Technical problems to be solved

[0009] To solve the problem that the existing algorithm uses a single-core processor as the computing platform, and when the processor is in the replacement link of the auction algorithm, each replacement needs to read from the memory through the bus, and each bus read has to go through operations such as fetching, decoding, executing, peripheral response, and storing to memory, and the program can only be executed sequentially without finding its parallelism, which greatly reduces the performance of the algorithm and thus affects its real-time performance. A coprocessor and a custom instruction design method for the replacement process in the data fusion algorithm based on the auction algorithm are provided.

[0010] (2) Technical solutions

[0011] A coprocessor for the replacement process in the data fusion algorithm based on the auction algorithm includes a coprocessor. The internal circuit of the coprocessor contains a control logic and a computing unit. The control logic unit is used to control the data flow of the entire computing unit and the data reading and writing with the storage cache. The computing unit is responsible for calculating a pair of data transmitted from the custom instruction, comparing whether the values read from two storage caches are the same, and sending the comparison result to the control logic.

[0012] As a preferred technical solution, it includes the following steps:

[0013] s1. After the main processor sends a custom instruction to the coprocessor module, the coprocessor saves the corresponding new pair of serial numbers in the instruction of the main processor in the register.

[0014] s2. The control logic starts from the zero-th number and reads and saves a pair of serial numbers from two storage caches, compares this pair of serial numbers with the serial numbers in the register, and in the next cycle, the control logic fetches the next group of serial numbers to continue the comparison.

[0015] s3. If the serial numbers in a certain group are the same as the serial numbers in the register, if they are the serial numbers in group K, then replace the serial numbers in group J in the storage and return the replaced serial numbers to the main processor for re-auction through the ROCC interface; if they are the serial numbers in group J, then directly replace the serial numbers in group K without returning the replaced serial numbers to the main processor. If all groups in the storage cache have been compared and there are no identical ones, then store this group of serial numbers in the register in the storage cache and increment the register value that records the number of groups in the record cache by one.

[0016] (3) Beneficial effects

[0017] The beneficial effects of the present invention are as follows:

[0018] 1. For an algorithm with repetitive and regular operation steps, implementing it with a finite state machine saves a large amount of time compared to implementing it with a processor. At the same time, through reasonable design, the two data retrievals and comparisons for each group of numbers are synchronized, optimizing at the algorithm level and doubling the efficiency.

[0019] 2. Utilizing the characteristics of the dual-port cache to achieve a direct interaction method between the peripheral device and the coprocessor, rather than accessing the on-chip storage through the ROCC interface. This eliminates the process of the main processor transporting data and also enables the coprocessor to access two numbers simultaneously instead of one number in the ROCC interface.

[0020] 3. Designing a coprocessor circuit for the replacement module in the algorithm and using the characteristics of the dual-port RAM to design a unique data interaction method between the bus peripheral device and the coprocessor to maximize the computing efficiency.

[0021] 4. Utilizing the built-in ROCC interface of the open-source processor RocketChip to control the behavior of the coprocessor through custom instructions to achieve the interaction between the coprocessor and the main processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 It is a schematic diagram of the coprocessor, processor, and dual-port storage peripheral;

[0024] Figure 2 It is the internal circuit diagram of the coprocessor module; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] In combination with the drawings, a further description is made of a coprocessor and a custom instruction design method for the replacement process in a data fusion algorithm based on an auction algorithm of the present invention. The following further details the present invention in combination with embodiments:

[0026] A coprocessor for the replacement process in a data fusion algorithm based on an auction algorithm, including a coprocessor. The internal circuit of the coprocessor contains control logic and a computing unit. The control logic unit is used to control the data flow of the entire computing unit and the data reading and writing with the storage cache. The computing unit is responsible for calculating a pair of data transmitted from a custom instruction, comparing whether the values read from two storage caches are the same, and sending the comparison result to the control logic.

[0027] Further, it includes the following steps:

[0028] S1. After the main processor sends a custom instruction to the coprocessor module, the coprocessor saves the corresponding new pair of serial numbers in the instruction of the main processor in a register.

[0029] S2. The control logic starts from the zero-th number and reads and saves a pair of serial numbers from two storage caches, compares this pair of serial numbers with the serial numbers in the register, and at the same time, in the next cycle, the control logic fetches the next pair of serial numbers to continue the comparison.

[0030] S3. If the serial numbers in a certain group are the same as those in the register, if they are the serial numbers in group K, then replace the serial numbers in group J in the storage, and return the replaced serial numbers to the main processor through the ROCC interface for re-auction; if they are the serial numbers in group J, then directly replace the serial numbers in group K, and do not return the replaced serial numbers to the main processor. If all groups in the storage cache have been compared and there are no identical ones, then store this group of serial numbers in the register in the storage cache and increment the register value that records the number of groups in the record cache by one.

[0031] Working principle: As Figure 1 shown, the coupling method of this coprocessor with the peripheral dual-port cache and the main processor is that the main processor and the coprocessor are connected through the ROCC interface of the coprocessor interface of RocketChip. Port A of the dual-port memory is connected to the TileLink bus, and port B of the dual-port memory is directly coupled to the coprocessor.

[0032] Table 1 shows the format of the RocketChip custom instruction, its operation code is custom0, the function code is 7’b0000000, and the source registers Rs1 and Rs2 and the destination register Rd are all valid.

[0033] Table 2 represents the formats of the source registers Rs1 and Rs2 and the destination register Rd, where the high 32 bits of each register are all 0, and the low 32 bits respectively represent specific serial numbers.

[0034] As Figure 2As shown, the internal circuit diagram of the coprocessor includes control logic, address generation logic, J sequence comparison logic, and K sequence comparison logic. After the processor sends a sequence number replacement instruction to the coprocessor, the control logic stores the values in the source register of the instruction into the J sequence register and the K sequence register, and at the same time starts sending control signals to the address generation logic to traverse all the data stored in the dual-port buffer from zero. In each cycle, the control logic sends an address to the two dual-port buffers, and two sequence number data are returned in the next cycle. These two sequence number data are respectively stored in the memory 1 sequence register and the memory 2 sequence register, and compared with the values in the J sequence register and the K sequence register respectively. If the comparison result is that the values in the K register are the same, in the next cycle, the value in the J register will be written back to the corresponding memory 1 at this address, and the processor will be notified through the Core Resp signal of the ROCC interface, and the value in the memory J will be written back to the processor; if the values in the J sequence register are the same, then in the next cycle, the control register will write the value in the K register back to the corresponding address in memory 2, and then notify the processor through Core Resp, and write the value -1 back to the processor. If no matching value is found after traversing all the numbers in the memory, the control logic will increment the value of the group number register and write the numbers in the J and K registers into memory 1 and memory 2 respectively with this address, then notify the processor through Core Resp, and send the number -1 back to the processor.

[0035] Table 1 shows the encoding format of the custom instruction

[0036]

[0037] Table 2 shows the contents of the source register and the destination register

[0038]

[0039]

[0040] The above embodiments only describe the preferred implementation manners of the present invention, and do not limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various variations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention should all fall within the protection scope of the present invention. The technical content claimed by the present invention has been fully recorded in the claims.

Claims

1. A method for designing custom instructions of a coprocessor in the replacement process of a data fusion algorithm based on an auction algorithm, characterized in that: Including the following steps: S1. After the main processor sends a custom instruction to the coprocessor module, the coprocessor saves a new pair of serial numbers corresponding to the instruction of the main processor in the register; S2. The control logic starts from the zero-th number, reads and saves a pair of serial numbers from the two storage caches, compares this pair of serial numbers with the serial numbers in the register, and in the next cycle, the control logic fetches the next pair of serial numbers to continue the comparison; S3. If the serial numbers in a certain pair are the same as those in the register, if they are the serial numbers in group K, then replace the serial numbers in group J in the storage, and return the replaced serial numbers to the main processor through the ROCC interface for re-auction; if they are the serial numbers in group J, then directly replace the serial numbers in group K and do not return the replaced serial numbers to the main processor; if all groups in the storage cache have been compared and there are no identical ones, then store this pair of serial numbers in the register in the storage cache and increment the register value that records the number of groups in the record cache by one; The internal of the coprocessor circuit includes a control logic and a computing unit. The control logic unit is used to control the data flow control of the entire computing unit and the data reading and writing with the storage cache. The computing unit is responsible for calculating whether a pair of data transmitted from the custom instruction and the values read from the two storage caches are the same, and sending the comparison result to the control logic.

Citation Information

Patent Citations

  • Delivery time period distribution method and device, equipment and storage medium

    CN113469590A

  • Method and apparatus for vectorized resource scheduling in distributed computing systems using tensors

    WO2021051772A1