Post-processing acceleration apparatus and method for object detection, data processing system

CN116263997BActive Publication Date: 2026-09-29SHANGHAI FUDAN MICROELECTRONICS GROUP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111520635.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-13
Publication Date
2026-09-29
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

经大量统计,对于SSD网络,进行后处理的95%以上的数据都为无效数据(也就是背景框),这大大浪费了后处理的时间,进而导致整个SSD网络检测时间增加

Benefits of technology

[0039]本发明实施例提供的用于目标检测的后处理加速装置及方法,在主机需要对存储在加速器端存储器中的数据进行后续处理时,不同于现有技术从加速器端存储器中读取数据后直接调用CPU去进行后处理,而是先通过硬件完成大部分无效数据的筛选,得到有效框数据,并将有效框数据的位置坐标信息输出,从而可使CPU处理数据量大大减少,大大减少后处理的时间,使SSD类型后处理的网络检测速度显著提高。而且,由于从加速器存储器传输到主机存储器的数据大大减少,因而有效节省了接口带宽。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116263997B_ABST
    Figure CN116263997B_ABST
Patent Text Reader

Abstract

The application provides a post-processing acceleration device and method for target detection and a data processing system. The device comprises: a start control module, configured to receive a start signal sent by a host and send configuration parameters set by the host to a direct storage access module after receiving the start signal; the direct storage access module, configured to perform addressing according to the configuration parameters, send a read data request to an accelerator end memory, and cache coordinate information corresponding to each address as frame data output to a comparison module one by one with the data returned by reading; the comparison module, configured to filter data in an anchor frame to obtain effective frame data, wherein one anchor frame comprises a plurality of frame data; and a format conversion module, configured to perform format conversion on the effective frame data and transmit the data after format conversion to the host. According to the application, the post-processing time can be reduced, and the interface bandwidth can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a post-processing acceleration device and method for target detection, and also to a data processing system. Background Technology

[0002] Typically, data preprocessed by the neural network passes through an AI accelerator, and the results are stored in the accelerator's memory. When the host computer receives a signal indicating that the computation is complete, it reads the results from the accelerator's memory and calls the CPU for post-processing. Post-processing operations generally performed on the CPU include: data layout format conversion, data type conversion, and comparison and filtering to obtain valid bounding boxes, thereby completing object detection. Its specific structure and process are as follows: Figure 1 and Figure 2 As shown.

[0003] SSD (Single Shot MultiBox Detector) is a one-stage general-purpose object detection algorithm. SSD uses a CNN (Convolutional Neural Network) to predict a series of bounding boxes and their corresponding object categories, and then uses the NMS (Non-maximum suppression) algorithm to obtain the final detection result. Generally, SSD post-processing data is arranged with an anchor box as the basic unit, and its data format is as follows: Figure 3 As shown, each anchor box includes its coordinates (x, y, w, h), background confidence data (p0), and confidence data for each category (p1, p2…pn). The key feature is finding the maximum value of p0-pn; if the maximum value is p0, it's a background box; otherwise, it's a potential target box (requiring further evaluation). Extensive statistical analysis shows that for SSD networks, over 95% of the data undergoing post-processing is invalid (i.e., background boxes), significantly wasting post-processing time and increasing the overall SSD network detection time. Furthermore, current solutions transfer all data from the AI ​​accelerator's memory to the host memory, resulting in wasted interface bandwidth and reduced accelerator efficiency. Summary of the Invention

[0004] One aspect of this invention provides a post-processing acceleration device and method for target detection, which reduces post-processing time and saves interface bandwidth.

[0005] Another aspect of the present invention provides a data processing system to improve data processing efficiency and save processing time.

[0006] To address the aforementioned technical problems, the embodiments of the present invention provide the following technical solutions:

[0007] On one hand, embodiments of the present invention provide a post-processing acceleration device for target detection, the device comprising:

[0008] The startup control module is used to receive a startup signal sent by the host, and after receiving the startup signal, send the configuration parameters set by the host to the direct storage access module.

[0009] The direct storage access module is used to address according to the configuration parameters, send a read data request to the accelerator-side memory, cache the coordinate information corresponding to each address, and output the frame data as a one-to-one correspondence with the read returned data to the comparison module.

[0010] The comparison module is used to filter the data in an anchor box to obtain valid frame data. An anchor box includes multiple frames of data.

[0011] The format conversion module is used to convert the format of the valid frame data and transmit the converted data to the host.

[0012] Optionally, the startup control module includes a startup register and a parameter register; the startup register is used to write the startup signal; and the parameter register is used to write the configuration parameters.

[0013] Optionally, the configuration parameters include: feature map parameters and anchor frame parameters.

[0014] Optionally, the comparison module inputs multi-path frame data in parallel.

[0015] Optionally, the comparison module includes: one or more binary comparators and a judgment unit; the binary comparator is used to compare the multi-path box data input in parallel in this round to obtain a comparison result; the judgment unit is used to determine whether the box data is background box data according to the comparison result, and if not, output the box data as valid box data to the format conversion module.

[0016] Optionally, each data frame has an independent enable bit.

[0017] Optionally, the comparison module further includes: a cache unit for storing the comparison result obtained in the previous round; and a memory comparison unit for comparing the comparison result obtained in the current round with the comparison result obtained in the previous round stored in the cache unit to obtain the final comparison result.

[0018] On the other hand, embodiments of the present invention also provide a post-processing acceleration method for target detection, the method comprising:

[0019] Upon receiving the start signal from the host, the system performs addressing according to the configuration parameters set by the host and caches the coordinate information corresponding to each address.

[0020] The data corresponding to the address is obtained by reading the accelerator end memory, and the cached address and the data are matched one by one as frame data;

[0021] Filter the data in an anchor box to obtain valid frame data; an anchor box includes multiple frames of data.

[0022] The data in the valid frames is formatted and then transmitted to the host.

[0023] Optionally, the method further includes: receiving the startup signal through a startup register; and obtaining the configuration parameters through a parameter register.

[0024] Optionally, filtering the data in an anchor box to obtain valid frame data includes:

[0025] Multiple binary comparators are used to simultaneously compare the multi-path frame data in one anchor frame of the current round to obtain the comparison result;

[0026] Based on the comparison result, determine whether the box data is background box data. If not, then the box data is considered valid box data.

[0027] Optionally, each data frame has an independent enable bit.

[0028] Optionally, filtering the data in an anchor box to obtain valid frame data further includes:

[0029] Cache the comparison results obtained in the previous round;

[0030] The comparison result obtained in this round is compared with the comparison result obtained in the previous round from the cache to obtain the final comparison result.

[0031] On the other hand, embodiments of the present invention also provide a data processing system, the system comprising: a host, an accelerator, a CPU, and a post-processing acceleration device for target detection as described above;

[0032] The host is used to send a start signal to the post-processing acceleration device for target detection and to set configuration parameters;

[0033] The post-processing acceleration device for target detection is used to receive a start signal sent by the host, and after receiving the start signal,

[0034] Addressing is performed according to the configuration parameters set by the host, data is obtained from the accelerator-side memory, the data is filtered, and valid frame data in a specific format is output to the host.

[0035] The host is also used to invoke the CPU to perform target detection on the valid frame data.

[0036] On the other hand, embodiments of the present invention also provide a chip including the post-processing acceleration device for target detection as described above.

[0037] On the other hand, embodiments of the present invention also provide a computer-readable storage medium, which is a non-volatile storage medium or a non-transient storage medium, on which a computer program is stored, and the computer program is executed by a processor to cause the above-described method to be performed.

[0038] On the other hand, embodiments of the present invention also provide a post-processing acceleration device for target detection, including a memory and a processor, wherein the memory stores source data that can be processed on the processor, and the processor processes the source data such that the above is executed.

[0039] The post-processing acceleration device and method for target detection provided in this invention differs from existing technologies that directly call the CPU for post-processing after reading data from the accelerator-side memory. Instead, when the host needs to perform post-processing on data stored in the accelerator-side memory, it first filters out most of the invalid data through hardware to obtain valid frame data and outputs the position coordinate information of the valid frame data. This significantly reduces the amount of data processed by the CPU, greatly reduces the post-processing time, and significantly improves the network detection speed of SSD-type post-processing. Moreover, since the amount of data transferred from the accelerator memory to the host memory is greatly reduced, interface bandwidth is effectively saved.

[0040] Accordingly, the data processing system provided in this embodiment of the invention can greatly improve data processing efficiency and save processing time. Attached Figure Description

[0041] Figure 1 This is a schematic diagram of the structure of post-processing technology in the prior art;

[0042] Figure 2 This is a flowchart of post-processing techniques in the prior art;

[0043] Figure 3 This is a schematic diagram of the data format of an anchor frame in existing SSD network post-processing;

[0044] Figure 4 This is a schematic diagram of a post-processing acceleration device for target detection according to an embodiment of the present invention;

[0045] Figure 5 This is a schematic diagram of a comparison module in a post-processing acceleration device for target detection according to an embodiment of the present invention;

[0046] Figure 6 This is a schematic diagram illustrating how a binary comparator performs multiple rounds of comparisons on data within an anchor frame, according to an embodiment of the present invention.

[0047] Figure 7 This is a schematic diagram of the data processing system according to an embodiment of the present invention;

[0048] Figure 8 yes Figure 7 A schematic diagram of data transmission in the data processing system of the embodiment shown;

[0049] Figure 9 This is a flowchart of a post-processing acceleration method for target detection according to an embodiment of the present invention. Detailed Implementation

[0050] To make the above-mentioned objectives, features and beneficial effects of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0051] To address the problem in existing SSD-based post-processing solutions that require transferring all data from the AI ​​accelerator's memory to the host memory, resulting in wasted interface bandwidth and reduced accelerator efficiency, this invention provides a post-processing acceleration device and method for target detection. Instead of directly transmitting data to the host after reading it from the accelerator's memory, this method first filters the data, removing invalid bounding box data to obtain valid bounding box data, and only transmits the valid bounding box data to the host.

[0052] like Figure 4 The diagram shown is a structural schematic of a post-processing acceleration device for target detection according to an embodiment of the present invention.

[0053] The post-processing acceleration device for target detection includes: a startup control module 401, a direct storage access module 402, a comparison module 403, and a format conversion module 404. Wherein:

[0054] The startup control module 401 is used to receive a startup signal sent by the host, and after receiving the startup signal, to send the configuration parameters set by the host to the direct storage access module 402.

[0055] The direct storage access module 402 is used to address according to the configuration parameters, send a read data request to the accelerator-side memory, cache the coordinate information corresponding to each address, and output the frame data as a one-to-one correspondence with the read returned data to the comparison module.

[0056] The comparison module 403 is used to filter the data in an anchor frame to obtain valid frame data. An anchor frame includes multiple frames of data.

[0057] The format conversion module 404 is used to convert the format of the valid frame data and transmit the converted data to the host.

[0058] In one non-limiting embodiment, a startup register and a parameter register can be configured in the startup control module 401. The host writes a startup signal to the startup register and configuration parameters to the parameter register.

[0059] Accordingly, the startup control module 401 can obtain the startup signal by reading the startup register and obtain the host configuration parameters by reading the parameter register. The configuration parameters mainly include: feature map parameters and anchor frame parameters. The feature map parameters include the length, width, and height of the feature map being compared. The anchor frame parameters include the length of each anchor frame, the effective category position in each anchor frame, the base address of the read accelerator memory, and the base address of the write processing system memory, etc.

[0060] The direct storage access module 402 addresses the accelerator memory according to the corresponding parameters in the order of an anchor frame, sends a read data request to the accelerator memory, and caches the coordinate information corresponding to each address, outputting it to the comparison module 403 in a one-to-one correspondence with the read return data.

[0061] In this embodiment of the invention, the comparison module 403 mainly compares the data of each box to determine the class with the highest confidence in each box. If the obtained class index value is 0, it represents a background box, and the data of that group is filtered out. If it is not a background box, the data of that group and the corresponding info information (w, h, n, max, idx, which represent width, height, number of groups, maximum number obtained by the comparator, and index value, respectively) are output to the format conversion module 404.

[0062] In one non-limiting embodiment, the comparison module 403 can simultaneously input multiple data that need to be compared, that is, input multiple box data in parallel. The number N of box data input in parallel depends on the circuit design. The higher the degree of parallelism, the faster the processing speed, but the more resources are consumed.

[0063] like Figure 5 The diagram shown is a structural schematic of a comparison module in a post-processing acceleration device for target detection according to an embodiment of the present invention.

[0064] In this embodiment, the comparison module 403 includes: one or more binary comparators 431, and a judgment unit 432. Wherein:

[0065] The binary comparator 431 is used to compare the multi-frame data input in parallel in this round and obtain the comparison result;

[0066] The judgment unit 432 is used to determine whether the box data is background box data based on the comparison result. If not, the box data is output as valid box data to the format conversion module.

[0067] In this embodiment of the invention, multiple binary comparators are designed in a pipelined manner, such as... Figure 6 As shown, the results of the second round of comparison will be available in the next clock cycle after the first round of comparison, and so on, until the largest number among all the data is finally found.

[0068] Considering that in some cases the number N of parallel input frame data may be limited by circuit capabilities, and the number of frame data to be compared is greater than N, multiple rounds of comparison are required. To address this, an independent enable bit can be set in each frame data channel; for example, the valid bit in each anchor frame can be determined based on the configured parameter register. Accordingly, when the number of parallel inputs to be compared is less than the number N, the enable bits of other input channels can be directly turned off; when the number of parallel inputs to be compared is greater than the number N, multiple rounds of comparison can be performed.

[0069] Therefore, in another non-limiting embodiment of the device of the present invention, such as Figure 5 and Figure 6 As shown, the comparison module 403 may further include a cache unit 433 and a memory comparison unit 434. The cache unit 433 is used to store the comparison result obtained in the previous round; the memory comparison unit 434 is used to compare the comparison result obtained in the current round with the comparison result obtained in the previous round stored in the cache unit 433 to obtain the final comparison result.

[0070] As can be seen, in this embodiment of the invention, the comparator is designed to be flexible, efficient, and capable of fulfilling the function of information following data transmission. After a fixed clock cycle, it can compare the largest number in an anchor frame and output its coordinates and other information. Its advantages are particularly pronounced when comparing a large amount of data, allowing for rapid comparison of the largest number and its position among any number of data points, while simultaneously outputting the coordinates of the anchor frame. The memory comparison unit can compare the larger of the results from two consecutive rounds, making it suitable for scenarios with a parallelism greater than N. Furthermore, it supports an arbitrary number of channels through channel enabling.

[0071] The format conversion module 404 receives the output of the comparison module 403, integrates the valid frame data into the standard format required for subsequent processing, and outputs it to the host. The specific conversion format is not limited in this embodiment of the invention. It can be converted according to the actual application system needs. For example, 64 bits of info information can be added before the 512 bits of valid frame data (i.e., non-background anchor frame data).

[0072] The post-processing acceleration device for target detection provided in this invention differs from existing technologies that directly call the CPU for post-processing after reading data from the accelerator-side memory. Instead, it first filters out most of the invalid data through hardware to obtain valid frame data and outputs the position coordinate information of the valid frame data. This significantly reduces the amount of data processed by the CPU, greatly reduces the post-processing time, and significantly improves the network detection speed of SSD-type post-processing. Moreover, since the amount of data transferred from the accelerator memory to the host memory is greatly reduced, interface bandwidth is effectively saved.

[0073] Accordingly, embodiments of the present invention also provide a data processing system, such as... Figure 7 As shown, the data processing system includes: a host 71, an accelerator 72, a CPU 74, and the aforementioned post-processing acceleration device 73 for target detection. Wherein:

[0074] The host 71 is used to send a start signal to the post-processing acceleration device for target detection and to set configuration parameters;

[0075] The post-processing acceleration device 73 for target detection is used to receive a start signal sent by the host 71, and after receiving the start signal, it addresses according to the configuration parameters set by the host 71, retrieves data from the memory at the end of the accelerator 72, filters the data, and outputs valid frame data in a specific format to the host 71.

[0076] Accordingly, the host 71 is also used to call the CPU 74 to perform target detection on the valid frame data.

[0077] It should be noted that in practical applications, the post-processing acceleration device 73 can be independent of the host 71 or as part of the host 71. This embodiment of the invention does not limit this.

[0078] Figure 8 It shows Figure 7 The data processing system shown illustrates how it processes the data after it has been computed by the artificial intelligence accelerator.

[0079] Verification shows that, using the solution of this invention, more than 95% of invalid data can be filtered out at the hardware level, which drastically reduces the amount of transmitted data and greatly saves interface bandwidth.

[0080] The acceleration scheme designed in this invention has a wide range of applications. It can configure the corresponding registers in the startup control unit according to the classification requirements in different scenarios. That is, in order to meet the data bandwidth requirements, sometimes some virtual data will be filled in an anchor box. This virtual data is actually invalid data. For this reason, an anchor box parameter register can be configured to represent which data in the anchor box is valid, that is, how many valid classes there are. Then, the enable bits of each input channel of the comparator can be controlled to eliminate the interference of invalid data and complete the filtering of valid data.

[0081] Accordingly, embodiments of the present invention also provide a post-processing acceleration method for target detection, as shown in Figure 9, which is a flowchart of the method, including the following steps:

[0082] Step 901: After receiving the start signal sent by the host, addressing is performed according to the configuration parameters set by the host, and the coordinate information corresponding to each address is cached.

[0083] Specifically, the startup signal can be received through the startup register; the configuration parameters can be obtained through the parameter register.

[0084] Step 902: Obtain the data corresponding to the address by reading the accelerator-side memory, and use the cached address and the data to form a frame data.

[0085] Step 903: Filter the data in an anchor frame to obtain valid frame data. An anchor frame includes multiple frames of data.

[0086] Specifically, multiple binary comparators can be used to simultaneously compare the multi-way box data in an anchor frame of the current round to obtain the comparison result; based on the comparison result, it is determined whether the box data is background box data; if not, the box data is taken as valid box data.

[0087] Furthermore, an independent enable bit can be set in each data frame to flexibly select the data to be compared in each round based on the number of parallel input data paths. In addition, when multiple rounds of comparison are required to complete the comparison of all data in an anchor frame, the comparison result obtained in the previous round can be cached; the comparison result obtained in the current round is compared with the cached comparison result obtained in the previous round to obtain the final comparison result.

[0088] Step 904: Convert the format of the valid frame data and transmit the converted data to the host.

[0089] In specific implementations, the aforementioned post-processing acceleration device for target detection can correspond to a chip with a corresponding function in network equipment and / or terminal equipment, such as a SOC (System-On-a-Chip), baseband chip, chip module, etc.

[0090] In specific implementation, the modules / units included in the various devices and products described in the above embodiments can be software modules / units, hardware modules / units, or a combination of both.

[0091] For example, for various devices and products applied to or integrated into a chip, each module / unit can be implemented using hardware methods such as circuits, or at least some modules / units can be implemented using software programs that run on a processor integrated within the chip, while the remaining (if any) modules / units can be implemented using hardware methods such as circuits; for various devices and products applied to or integrated into a chip module, each module / unit can be implemented using hardware methods such as circuits, and different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components of the chip module, or at least some modules / units can be implemented using hardware methods such as circuits. The components can be implemented using software programs that run on the processor integrated within the chip module. The remaining (if any) modules / units can be implemented using hardware methods such as circuits. For various devices and products applied to or integrated into the terminal, each of its components / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or in different components within the terminal. Alternatively, at least some modules / units can be implemented using software programs that run on the processor integrated within the terminal, while the remaining (if any) modules / units can be implemented using hardware methods such as circuits.

[0092] This invention also provides a computer-readable storage medium, which is a non-volatile or non-transient storage medium, storing a computer program thereon. When the computer program is run by a processor, it executes the steps in the above-described method embodiments.

[0093] This invention also provides a post-processing acceleration device for target detection, including a memory and a processor. The memory stores source data that needs to be processed on the processor, and the processor executes the steps in the above-described method embodiments when processing the source data.

[0094] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article indicates that the preceding and following related objects have an "or" relationship.

[0095] In the embodiments of this invention, "multiple" refers to two or more.

[0096] The descriptions of "first," "second," etc., appearing in the embodiments of this invention are for illustrative purposes and to distinguish the objects being described. They do not indicate any particular order and do not imply any special limitation on the number of devices in the embodiments of this invention. They do not constitute any limitation on the embodiments of this invention.

[0097] The embodiments provided by this invention can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. It should be understood that in the various embodiments of this invention, the sequence number of the above processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this invention.

[0098] In the several embodiments provided by this invention, it should be understood that the disclosed methods, apparatuses, and systems can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and other division methods may exist in actual implementation; for example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0099] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0100] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can be physically arranged separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0101] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.

Claims

1. A post-processing acceleration device for target detection, characterized in that, The device includes: A startup control module is provided, which includes a startup register and a parameter register. The startup register is used to write the startup signal sent by the host. The parameter register is used to write the configuration parameters set by the host. The configuration parameters include feature map parameters and anchor frame parameters. The startup control module receives the startup signal by reading the startup register, and after receiving the startup signal, reads the configuration parameters from the parameter register and sends the configuration parameters to the direct memory access module. The direct storage access module is used to address according to the anchor frame parameters in the configuration parameters, send a read data request to the accelerator-side memory, cache the coordinate information corresponding to each address, and output the frame data as a one-to-one correspondence with the read returned data to the comparison module. The comparison module is used to filter the data in an anchor frame to obtain valid frame data. An anchor frame includes multiple frames. During the filtering process, multiple sets of frame data are input in parallel, and the multiple sets of frame data are compared to filter out background frame data to obtain valid frame data. The format conversion module is used to convert the format of the valid frame data and transmit the converted data to the host.

2. The apparatus according to claim 1, characterized in that, The comparison module includes: one or more binary comparators, and a judgment unit; The binary comparator is used to compare the multi-frame data input in parallel in this round and obtain the comparison result; The judgment unit is used to determine whether the box data is background box data based on the comparison result. If not, the box data is output as valid box data to the format conversion module.

3. The apparatus according to claim 2, characterized in that, Each data frame has an independent enable bit.

4. The apparatus according to claim 3, characterized in that, The comparison module further includes: A cache unit is used to store the comparison results obtained in the previous round; The memory comparison unit is used to compare the comparison result obtained in the current round with the comparison result obtained in the previous round stored in the cache unit to obtain the final comparison result.

5. A post-processing acceleration method for target detection, characterized in that, The method includes: After receiving the startup signal sent by the host through the startup register, the configuration parameters set by the host are obtained by reading the parameter register. The configuration parameters include feature map parameters and anchor box parameters. Addressing is performed according to the anchor box parameters in the configuration parameters, and the coordinate information corresponding to each address is cached. The data corresponding to the address is obtained by reading the accelerator-side memory, and the cached address and the data are matched one-to-one as frame data; The data in an anchor frame is filtered to obtain valid frame data. An anchor frame includes multiple frames. During the filtering process, the data from multiple frames are compared in parallel to filter out background frames and obtain valid frame data. The data in the valid frames is formatted and then transmitted to the host.

6. The method according to claim 5, characterized in that, The filtering of data within an anchor frame to obtain valid frame data includes: Multiple binary comparators are used to simultaneously compare the multi-path frame data in one anchor frame of the current round to obtain the comparison result; Based on the comparison result, determine whether the box data is background box data. If not, then the box data is considered valid box data.

7. The method according to claim 6, characterized in that, Each data frame has an independent enable bit.

8. The method according to claim 7, characterized in that, The process of filtering the data within an anchor frame to obtain valid frame data also includes: Cache the comparison results obtained in the previous round; The comparison result obtained in this round is compared with the comparison result obtained in the previous round from the cache to obtain the final comparison result.

9. A data processing system, characterized in that, The system includes: a host, an accelerator, a CPU, and a post-processing acceleration device for target detection as described in any one of claims 1 to 4; The host is used to send a start signal to the post-processing acceleration device for target detection and to set configuration parameters; The post-processing acceleration device for target detection is used to receive a start signal sent by the host, and after receiving the start signal, Addressing is performed according to the configuration parameters set by the host, data is obtained from the accelerator-side memory, the data is filtered, and valid frame data in a specific format is output to the host. The host is also used to invoke the CPU to perform target detection on the valid frame data.

10. A chip, characterized in that, Includes the post-processing acceleration device for target detection as described in any one of claims 1 to 4.

11. A computer-readable storage medium, said computer-readable storage medium being a non-volatile storage medium or a non-transient storage medium, having stored thereon a computer program, characterized in that, When the computer program is run by the processor, the method described in any one of claims 5 to 8 is performed.

12. A post-processing acceleration device for target detection, comprising a memory and a processor, wherein the memory stores source data to be processed on the processor, characterized in that, The processor performs the method according to any one of claims 5 to 8 when processing source data.

Citation Information

Patent Citations

  • Image target detection method and system, electronic equipment and storage medium

    CN110781819A