A neural network acceleration device and method for pedestrian re-identification

By designing a feature cache module for multi-branch network structures and a dense convolutional module for parallel computing in the neural network acceleration device, the problem of being unable to effectively accelerate the multi-branch structure of the all-round network in the prior art is solved, and the effect of reducing delay and power consumption is achieved.

CN114782890BActive Publication Date: 2025-06-10SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210371662.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-11
Publication Date
2025-06-10
Estimated Expiration
2042-04-11

AI Technical Summary

Technical Problem

Existing neural network accelerators cannot effectively accelerate multi-branch structures in all-around networks, resulting in increased delay and power consumption, and dense convolution and depth separable convolution cascades have caused difficulties in data storage and allocation.

Method used

A pedestrian re-identification neural network acceleration device is designed. By setting up a feature cache module with a multi-branch network structure, the cache of intermediate features is realized, and the delay and power consumption is reduced; and through parallel calculations of the dense convolution module and the deep separable convolution module, the number and number of data written back to the cache module are reduced.

Benefits of technology

The intermediate feature cache of multi-branch network structure is realized, which reduces the delay and power consumption, and reduces the number of data write-backs through parallel calculations, further reducing the delay and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782890B_ABST
    Figure CN114782890B_ABST
Patent Text Reader

Abstract

The present invention discloses a neural network acceleration device and method for pedestrian re-identification. A neural network acceleration device for pedestrian re-identification includes: a control module, a cache module, a row cache module, a dense convolution module, a depthwise separable convolution module, and a storage module, wherein the cache module includes a weight cache module and a feature cache module, and the feature cache module is a multi-branch network structure. By setting the feature cache module with a multi-branch network structure, the input features obtained by calculation are used to obtain intermediate features and cached, realizing the caching of the intermediate features of the multi-branch network structure, without the need to frequently access the memory to obtain the intermediate features, reducing latency and power consumption; by setting the parallel calculation of the dense convolution module and the depthwise separable convolution module, the number and quantity of data written back to the cache module are reduced, thereby further reducing latency and power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of neural network computing, and in particular to a neural network acceleration device and method for pedestrian re-identification. Background Art

[0002] Pedestrian re-identification technology is an image retrieval technology that uses computer vision technology to match pictures of the same pedestrian taken by different cameras, and has broad application prospects in the field of intelligent security. When performing pedestrian re-identification, not only the global and holistic features of people are required, but also the local and fine-grained features, that is, the full-scale features. Based on this, researchers have proposed an Omni-Scale Network (OSNet) based on deep learning that can mine different-scale information of input pictures. This network structure uses multiple branches to obtain different-scale information, and finally fuses the different-scale information to improve the accuracy of pedestrian re-identification. In addition, the network also uses depthwise separable convolutions to reduce the number of parameters and the amount of computation. The Omni-Scale Network includes four branches, and the four branches are respectively composed of 1, 2, 3, and 4 Lite 3×3 convolutions in cascade. The intermediate calculations of the Omni-Scale Network need to retain data with multiple-scale information and frequently access memory, resulting in latency and additional power consumption. Existing neural network accelerators lack research on algorithm networks with multi-branch structures, so the existing accelerator architectures cannot be used to accelerate the Omni-Scale Network. In addition, the cascading of traditional dense convolutions and depthwise separable convolutions in the Omni-Scale Network causes difficulties in data storage arrangement. Summary of the Invention

[0003] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0004] To this end, an object of an embodiment of the present invention is to provide a neural network acceleration device and method for pedestrian re-identification, which realizes caching of intermediate features of a multi-branch network structure, reduces latency and power consumption, and realizes parallel operations of dense convolution calculation and depthwise separable convolution calculation.

[0005] In order to achieve the above technical objectives, the technical solutions adopted in the embodiments of the present invention include:

[0006] In a first aspect, an embodiment of the present invention provides a neural network acceleration device for pedestrian re-identification, including:

[0007] A control module for generating read control signals and write control signals;

[0008] The cache module includes a weight cache module and a feature cache module. The weight cache module is used to obtain and cache weights according to the write control signal. The feature cache module has a multi-branch network structure and is used to obtain input features according to the write control signal, calculate intermediate features based on the input features, and cache them;

[0009] The row cache module is used to read the weights and the intermediate features according to the read control signal, and rearrange the weights and the intermediate features to obtain a weight vector and a feature vector;

[0010] The dense convolution module is used to perform dense convolution calculations according to the weight vector and the feature vector to generate a first feature;

[0011] The depthwise separable convolution module is used to generate a second feature from the first feature, where the second feature is the depthwise separable convolution weight, and perform depthwise separable convolution calculations according to the first feature and the second feature to obtain a third feature;

[0012] The storage module is used to store the first feature and the third feature into the feature cache module.

[0013] In addition, a neural network acceleration device for pedestrian re-identification according to the above embodiments of the present invention may further have the following additional technical features:

[0014] Furthermore, in a neural network acceleration device for pedestrian re-identification according to an embodiment of the present invention, the multi-branch network structure of the feature cache module includes a first feature cache module, two second feature cache modules, three third feature cache modules, three fourth feature cache modules, and a fifth feature cache module;

[0015] The first feature cache module acquires the input features according to the write control signal and caches them, performs convolution calculation based on the input features to generate intermediate output features and caches them in the first of the second feature cache modules. The intermediate output features generate a first convolution result, a second convolution result, a third convolution result, and a fourth convolution result through Lite 3×3 convolution calculation, stores the first convolution result in the first of the third feature cache modules, stores the second convolution result in the first of the fourth feature cache modules, stores the third convolution result in the second of the fourth cache modules, stores the fourth convolution result in the third of the fourth cache modules. The first convolution result and the second convolution result are fused to generate a first fusion result and stored in the fifth feature cache module. The first fusion result and the third convolution result are fused to generate a second fusion result and stored in the second of the third feature cache modules. The second fusion result and the fourth convolution result are fused to generate a third fusion result and stored in the second of the second feature cache modules. Convolution calculation is performed based on the third fusion result to generate a fifth convolution result and stored in the third of the third feature cache modules. Residual fusion is performed based on the input features and the fifth convolution result to generate the intermediate features and stored in the first feature cache module.

[0016] Further, in an embodiment of the present invention, the first feature cache module and the fourth feature cache module adopt a ping-pong cache structure.

[0017] Further, in an embodiment of the present invention, the storage methods of the first feature cache module, the second feature cache module, the third feature cache module, the fourth feature cache module, and the fifth feature cache module include feature width first and feature channel first.

[0018] Further, in an embodiment of the present invention, the weight cache module adopts a ping-pong cache structure.

[0019] Further, in an embodiment of the present invention, the dense convolution module includes a multiplication module and an addition tree;

[0020] The multiplication module performs multiplication operation according to the weight vector and the feature vector to obtain a product vector; the addition tree adds the product vectors to obtain the first feature.

[0021] Further, in an embodiment of the present invention, the neural network acceleration device for pedestrian re-identification further includes a pooling module;

[0022] The pooling module performs max pooling calculation or average pooling calculation based on the first feature to obtain a pooling result; the storage module stores the pooling result in the feature cache module.

[0023] In a second aspect, an embodiment of the present invention provides a neural network acceleration method for pedestrian re-identification. The method is applied to a neural network acceleration device for pedestrian re-identification. The neural network acceleration device for pedestrian re-identification includes a control module, a cache module, a row cache module, a dense convolution module, a depthwise separable convolution module, and a storage module. The cache module includes a weight cache module and a feature cache module. The feature cache module has a multi-branch network structure. The method includes:

[0024] Generating a read control signal and a write control signal through the control module;

[0025] Obtaining and caching weights through the weight cache module according to the write control signal;

[0026] Obtaining input features through the feature cache module according to the write control signal, and calculating and caching intermediate features based on the input features;

[0027] Obtaining the weights and the intermediate features through the row cache module according to the read control signal, and rearranging the weights and the intermediate features to obtain a weight vector and a feature vector;

[0028] Performing dense convolution calculation through the dense convolution module according to the weight vector and the feature vector to generate a first feature;

[0029] Generating a second feature through the depthwise separable convolution module according to the first feature, where the second feature is a depthwise separable convolution weight;

[0030] Performing depthwise separable convolution calculation through the depthwise separable convolution module according to the first feature and the second feature to obtain a third feature;

[0031] Storing the first feature and the third feature in the feature cache module through the storage module.

[0032] Further, in an embodiment of the present invention, the multi-branch network structure of the feature cache module includes a first feature cache module, two second feature cache modules, three third feature cache modules, three fourth feature cache modules, and a fifth feature cache module;

[0033] The step of obtaining input features through the feature cache module according to the write control signal, and calculating and caching intermediate features based on the input features includes:

[0034] According to the writing control signal, obtain the input feature through the first feature cache module and cache it;

[0035] Perform convolution calculation according to the input feature, generate an intermediate output feature and cache it in the first of the second feature cache modules;

[0036] The intermediate output feature generates a first convolution result, a second convolution result, a third convolution result, and a fourth convolution result through Lite 3×3 convolution calculation, store the first convolution result in the first of the third feature cache modules, store the second convolution result in the first of the fourth feature cache modules, store the third convolution result in the second of the fourth cache modules, and store the fourth convolution result in the third of the fourth cache modules;

[0037] Fuse the first convolution result and the second convolution result to generate a first fusion result and store it in the fifth feature cache module;

[0038] Fuse the first fusion result and the third convolution result to generate a second fusion result and store it in the second of the third feature cache modules;

[0039] Fuse the second fusion result and the fourth convolution result to generate a third fusion result and store it in the second of the second feature cache modules;

[0040] Perform convolution calculation according to the third fusion result to generate a fifth convolution result and store it in the third of the third feature cache modules;

[0041] Perform residual fusion according to the input feature and the fifth convolution result to generate the intermediate feature and store it in the first feature cache module.

[0042] Further, in an embodiment of the present invention, the neural network acceleration device for pedestrian re-identification further includes a pooling module;

[0043] The neural network acceleration method for pedestrian re-identification further includes:

[0044] According to the first feature, perform max pooling calculation or average pooling calculation through the pooling module to obtain a pooling result;

[0045] Store the pooling result in the feature cache module through the storage module.

[0046] The advantages and beneficial effects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of this application:

[0047] In the embodiment of the present invention, by setting a feature cache module with a multi-branch network structure, intermediate features are calculated from the obtained input features and cached, realizing the caching of intermediate features of the multi-branch network structure. There is no need to frequently access the memory to obtain intermediate features, reducing latency and power consumption. By setting the parallel computing of the dense convolution module and the depthwise separable convolution module, the number and frequency of data written back to the cache module are reduced, thereby further reducing latency and power consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the related technical solution drawings in the embodiments of the present application or the prior art. It should be understood that the drawings below only facilitate the clear expression of some embodiments of the technical solutions in the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 It is a schematic structural diagram of a specific embodiment of a neural network acceleration device for pedestrian re-identification according to the present invention;

[0050] Figure 2 It is a schematic structural diagram of a weight cache module of a specific embodiment of a neural network acceleration device for pedestrian re-identification according to the present invention;

[0051] Figure 3 It is a schematic diagram of the storage method of a feature cache module of a specific embodiment of a neural network acceleration device for pedestrian re-identification according to the present invention;

[0052] Figure 4 It is a schematic diagram of inter-layer parallel operation of a specific embodiment of a neural network acceleration device for pedestrian re-identification according to the present invention;

[0053] Figure 5 It is a schematic flowchart of a specific embodiment of a neural network acceleration method for pedestrian re-identification according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application. For the step numbers in the following embodiments, they are only set for ease of explanation and illustration, and no limitation is placed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0055] In the description, claims and drawings of the present invention, terms such as "first", "second", "third" and "fourth" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices.

[0056] Reference to "embodiments" in the present invention means that specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0057] Person re-identification technology is an image retrieval technology that uses computer vision technology to match pictures of the same person taken by different cameras, and has broad application prospects in the field of intelligent security. When performing person re-identification, not only the global and holistic features of a person are required, but also the local and fine-grained features, that is, the full-scale features, need to be mined. Based on this, researchers have proposed an Omni-Scale Network (OSNet) based on deep learning that can mine different-scale information of input pictures. This network structure uses multiple branches to obtain different-scale information, and finally fuses the different-scale information to improve the accuracy of person re-identification. In addition, the network also uses depthwise separable convolutions to reduce the number of parameters and the amount of computation. The Omni-Scale Network includes four branches, and the four branches are respectively composed of 1, 2, 3, and 4 Lite 3×3 convolutions in cascade. The intermediate calculations of the Omni-Scale Network need to retain data with multiple-scale information and frequently access memory, resulting in latency and additional power consumption. Existing neural network accelerators lack research on algorithm networks with multi-branch structures, so the existing accelerator architectures cannot be used to accelerate the Omni-Scale Network. In addition, the cascading of traditional dense convolution and depthwise separable convolution in the Omni-Scale Network causes difficulties in data storage arrangement.

[0058] To this end, the present invention proposes a neural network acceleration device and method for pedestrian re-identification. By setting a feature caching module with a multi-branch network structure, intermediate features are calculated and cached from the obtained input features, realizing the caching of intermediate features of the multi-branch network structure. There is no need to frequently access memory to obtain intermediate features, reducing latency and power consumption. By setting the parallel computing of the dense convolution module and the depthwise separable convolution module, the number and frequency of data written back to the caching module are reduced, thereby further reducing latency and power consumption.

[0059] Next, a neural network acceleration device and method for pedestrian re-identification according to an embodiment of the present invention will be described in detail with reference to the accompanying drawings. First, a neural network acceleration device for pedestrian re-identification according to an embodiment of the present invention will be described with reference to the accompanying drawings.

[0060] Referring to Figure 1 , a neural network acceleration device for pedestrian re-identification in an embodiment of the present invention includes:

[0061] A control module for generating read control signals and write control signals;

[0062] A caching module including a weight caching module and a feature caching module. The weight caching module is used to obtain and cache weights according to the write control signal. The feature caching module has a multi-branch network structure and is used to obtain input features according to the write control signal, calculate intermediate features from the input features, and cache them;

[0063] A row caching module for reading the weights and the intermediate features according to the read control signal, and rearranging the weights and the intermediate features to obtain a weight vector and a feature vector;

[0064] A dense convolution module for performing dense convolution calculations according to the weight vector and the feature vector to generate a first feature;

[0065] A depthwise separable convolution module for generating a second feature, which is the depthwise separable convolution weight, from the first feature, and performing depthwise separable convolution calculations according to the first feature and the second feature to obtain a third feature;

[0066] A storage module for storing the first feature and the third feature into the feature caching module.

[0067] Among them, the signals of the external interface of a neural network acceleration device for pedestrian re-identification in an embodiment of the present invention include AXI4 protocol signals for external communication, an input clock clk signal, an input reset rst_n signal, a data_req (signal for fetching data externally), and a cal_over (signal for sending out data after calculation is completed).

[0068] By setting a feature cache module with a multi-branch network structure, the present invention calculates the obtained input features to obtain intermediate features and caches them, realizing the caching of the intermediate features of the multi-branch network structure, without the need to frequently access memory to obtain intermediate features, reducing latency and power consumption; by setting the parallel calculation of the dense convolution module and the depthwise separable convolution module, the number and frequency of data written back to the cache module are reduced, thereby further reducing latency and power consumption.

[0069] As an alternative implementation, the multi-branch network structure of the feature cache module includes a first feature cache module, two second feature cache modules, three third feature cache modules, three fourth feature cache modules, and a fifth feature cache module;

[0070] The first feature cache module obtains and caches the input feature according to the write control signal, performs convolution calculation on the input feature to generate an intermediate output feature and caches it in the first of the second feature cache modules. The intermediate output feature generates a first convolution result, a second convolution result, a third convolution result, and a fourth convolution result through Lite 3×3 convolution calculation, and stores the first convolution result in the first of the third feature cache modules, stores the second convolution result in the first of the fourth feature cache modules, stores the third convolution result in the second of the fourth cache modules, stores the fourth convolution result in the third of the fourth feature cache modules. The first convolution result and the second convolution result are fused to generate a first fusion result and stored in the fifth feature cache module. The first fusion result and the third convolution result are fused to generate a second fusion result and stored in the second of the third feature cache modules. The second fusion result and the fourth convolution result are fused to generate a third fusion result and stored in the second of the second feature cache modules. Convolution calculation is performed on the third fusion result to generate a fifth convolution result and stored in the third of the third feature cache modules. Residual fusion is performed on the input feature and the fifth convolution result to generate the intermediate feature and stored in the first feature cache module.

[0071] Specifically, the input feature passes through a 1×1 convolution and then performs residual fusion with the fifth convolution result to generate the intermediate feature and store it in the first feature cache module.

[0072] The design of the multi-branch network structure of the feature cache module in the embodiment of the present invention realizes the caching of the intermediate features of the multi-branch network structure, without the need to frequently access memory to obtain intermediate features, reducing latency and power consumption.

[0073] As an alternative implementation, the first feature cache module and the fourth feature cache module adopt a ping-pong cache structure.

[0074] Among them, the ping-pong cache structure of the first feature cache module and the fourth feature cache module includes feature cache block 0 and feature cache block 1. When one feature cache block is performing a feature reading operation, the other feature cache block can simultaneously perform a feature writing operation.

[0075] Specifically, in the embodiments of the present invention, when the first feature cache module and the fourth feature cache module adopting the ping-pong cache structure cache and read features, the features are cached into the corresponding feature cache block (feature cache block 0 or feature cache block 1) according to the cache control signal. When the corresponding feature cache block is full, the feature reading is executed, and at the same time, the features are continuously written into the other feature cache block, so that the time for obtaining features overlaps with the time for feature reading and calculation, improving the operation speed.

[0076] Refer to Figure 3 , as an optional implementation manner, the storage methods of the first feature cache module, the second feature cache module, the third feature cache module, the fourth feature cache module, and the fifth feature cache module include feature width-first and feature channel-first.

[0077] Among them, the feature width-first storage method is applicable to picture type data and input features of dense convolution with a convolution kernel greater than 1, facilitating the convolution window sliding and data reuse between adjacent convolution operations.

[0078] Specifically, in an embodiment of the present invention, a picture of size 256×128×3 is stored through the feature cache module. First, the data of the first row in the W (width) direction of the first channel of the picture is stored, then the data of the second row of the first channel is stored, and so on until the 128th data in the 128th row of the first channel is stored. According to the storage method of the first channel, the data of the second channel is stored until the entire picture is stored.

[0079] Feature channel-first is applicable to the storage of input features of dense convolution with a convolution kernel size of 1 to meet the data supply requirements when the input channels are large.

[0080] Specifically, in an embodiment of the present invention, the input of 16×8×512 channels of the feature cache module: first, the first data in the first row of 512 channels is stored, then the second data in the first row of 512 channels is stored until all the data is stored.

[0081] By setting the two storage methods of the feature cache module, the requirements of different convolutional layers for fast data supply are met, and the delay of reading data is reduced.

[0082] Refer to Figure 2, As an alternative embodiment, the weight cache module adopts a ping-pong cache structure.

[0083] Among them, the ping-pong cache structure of the weight cache module includes a weight cache block 0 and a weight cache block 1. When one weight cache block is performing a weight reading operation, the other weight cache block can simultaneously perform a weight writing operation.

[0084] Specifically, in the embodiment of the present invention, when the weight cache module adopting the ping-pong cache structure caches and reads weights, the weights are cached into the corresponding weight cache block (weight cache block 0 or weight cache block 1) according to the cache control signal. When the corresponding weight cache block is full, the weight reading is performed, and at the same time, the weights are continuously written into the other weight cache block, so as to overlap the time for obtaining weights with the time for weight reading and calculation, thereby improving the operation speed.

[0085] As an alternative embodiment, the dense convolution module includes a multiplication module and an addition tree;

[0086] The multiplication module performs a multiplication operation according to the weight vector and the feature vector to obtain a product vector; the addition tree adds the product vectors to obtain the first feature.

[0087] Among them, the multiplication module consists of multiple multipliers to form a multiplication array, and multiple multiplication calculations in traditional convolution are performed simultaneously; the addition tree adds the products (product vectors) output by the multiplication module to fuse features of different scales.

[0088] Specifically, referring to Figure 4 , in an embodiment of the present invention, taking the calculation process where the input feature size is 16×8×128, 1×1 dense convolution and 3×3 depthwise separable convolution layers are parallel, and the output feature size is still 16×8×128 as an example. The lengths of the feature vector and the weight vector are 512. The feature vector and the weight vector are input into the multiplication module for multiplication operation to obtain a product vector with a vector length of 512. After the product vectors are added by the addition tree, 4 features output by the 1×1 convolution are obtained, that is, the first feature (stored in the depthwise separable convolution module). The depthwise separable convolution module generates 3×3 depthwise separable convolution weights according to the first feature, that is, the second feature, and starts to perform depthwise separable convolution according to the first feature and the second feature to obtain an output feature, that is, the third feature. At this time, only the third feature needs to be written back to the cache module. During the inter-layer parallel operation process of the above-mentioned embodiment, the proportion of the number of features saved from being written back to the cache is 16×8×128 / (16×8×128 + 16×8×128) = 0.5, that is, it is saved by one time, achieving the purpose of accelerating the calculation speed and reducing the power consumption.

[0089] In an embodiment of the present invention, the neural network acceleration device for pedestrian re-identification further includes a pooling module;

[0090] The pooling module performs maximum pooling calculation or average pooling calculation according to the first feature to obtain a pooling result; the storage module stores the pooling result in the feature cache module.

[0091] The setting of the pooling module in the embodiment of the present invention realizes the inter-layer parallel processing of dense convolution and pooling calculation, reduces the number and quantity of data written back to the cache module, thereby further reducing the delay and power consumption.

[0092] In summary, the neural network acceleration device for pedestrian re-identification in the embodiment of the present invention realizes four acceleration modes of traditional dense convolution single operation, traditional dense convolution and depthwise separable convolution inter-layer parallel operation, traditional dense convolution and pooling inter-layer parallel operation, and feature fusion operation. The above four calculations cover all the calculations of the multi-branch neural network, and play the effect of accelerating the operation of the multi-branch neural network.

[0093] Secondly, referring to Figure 5 , the embodiment of the present invention proposes a neural network acceleration method for pedestrian re-identification. The method is applied to a neural network acceleration device for pedestrian re-identification. The neural network acceleration device for pedestrian re-identification includes a control module, a cache module, a row cache module, a dense convolution module, a depthwise separable convolution module and a storage module. The cache module includes a weight cache module and a feature cache module. The feature cache module is a multi-branch network structure. The method includes:

[0094] S101. Generate a read control signal and a write control signal through the control module;

[0095] S102. Obtain and cache weights through the weight cache module according to the write control signal;

[0096] S103. Obtain input features through the feature cache module according to the write control signal, and calculate and cache intermediate features according to the input features;

[0097] Wherein, the multi-branch network structure of the feature cache module includes a first feature cache module, two second feature cache modules, three third feature cache modules, three fourth feature cache modules and a fifth feature cache module;

[0098] Specifically, S103 can be further divided into the following steps S1031-S1038:

[0099] Step S1031: Obtain and cache the input feature through the first feature cache module according to the writing control signal;

[0100] Among them, the first cache module adopts a ping-pong cache structure.

[0101] Specifically, the ping-pong cache structure of the first feature cache module includes feature cache block 0 and feature cache block 1. When one feature cache block is performing a feature reading operation, the other feature cache block can simultaneously perform a feature writing operation.

[0102] Step S1032: Perform convolution calculation according to the input feature, generate an intermediate output feature, and cache it in the first second feature cache module;

[0103] Step S1033: The intermediate output feature generates a first convolution result, a second convolution result, a third convolution result, and a fourth convolution result through Lite 3×3 convolution calculation, store the first convolution result in the first third feature cache module, store the second convolution result in the first fourth feature cache module, store the third convolution result in the second fourth cache module, and store the fourth convolution result in the third fourth cache module;

[0104] Among them, the fourth cache module adopts a ping-pong cache structure.

[0105] Specifically, the ping-pong cache structure of the fourth feature cache module includes feature cache block 0 and feature cache block 1. When one feature cache block is performing a feature reading operation, the other feature cache block can simultaneously perform a feature writing operation.

[0106] Step S1034: Fuse the first convolution result and the second convolution result to generate a first fusion result and store it in the fifth feature cache module;

[0107] Step S1035: Fuse the first fusion result and the third convolution result to generate a second fusion result and store it in the second third feature cache module;

[0108] Step S1036: Fuse the second fusion result and the fourth convolution result to generate a third fusion result and store it in the second second feature cache module;

[0109] Step S1037: Perform convolution calculation according to the third fusion result to generate a fifth convolution result and store it in the third third feature cache module;

[0110] Step S1038: Perform residual fusion based on the input feature and the fifth convolution result to generate the intermediate feature and store it in the first feature cache module.

[0111] S104: According to the read control signal, obtain the weight and the intermediate feature through the row cache module, and rearrange the weight and the intermediate feature to obtain a weight vector and a feature vector;

[0112] S105: According to the weight vector and the feature vector, perform dense convolution calculation through the dense convolution module to generate a first feature;

[0113] S106: According to the first feature, generate a second feature through the depthwise separable convolution module;

[0114] Wherein, the second feature is the depthwise separable convolution weight.

[0115] S107: According to the first feature and the second feature, perform depthwise separable convolution calculation through the depthwise separable convolution module to obtain a third feature;

[0116] S108: Store the first feature and the third feature in the feature cache module through the storage module.

[0117] In an embodiment of the present invention, the neural network acceleration device for pedestrian re-identification further includes a pooling module;

[0118] The above-mentioned method for accelerating a neural network for pedestrian re-identification further includes:

[0119] Perform max pooling calculation or average pooling calculation on the first feature through the pooling module to obtain a pooling result;

[0120] Store the pooling result in the feature cache module through the storage module.

[0121] The content in the above method embodiments is applicable to the system embodiments. The functions specifically implemented by the system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0122] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially concurrently or the blocks can sometimes be executed in reverse order. Further, the embodiments presented and described in the flowcharts of the present application are provided by way of example for purposes of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are executed independently.

[0123] In addition, although the present application has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present application. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Accordingly, those of ordinary skill in the art will be able to implement the present application as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.

[0124] It should be understood that the various parts of the present application may be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, a plurality of steps or methods may be implemented by software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented in hardware, as in another embodiment, any one of the following techniques well known in the art or a combination thereof may be used: discrete logic circuits having logic gates for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0125] In the foregoing description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0126] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.

[0127] The above is a specific description of the preferred embodiment of the present application, but the present application is not limited to the described embodiment. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.

Claims

1. A neural network acceleration device for pedestrian re-identification, characterized in that, it includes: A control module for generating read control signals and write control signals; A cache module, including a weight cache module and a feature cache module. The weight cache module is used to obtain and cache weights according to the write control signal. The feature cache module is a multi-branch network structure, used to obtain input features according to the write control signal, and used to calculate and cache intermediate features according to the input features; A row cache module for reading the weights and the intermediate features according to the read control signal, and for rearranging the weights and the intermediate features to obtain a weight vector and a feature vector; A dense convolution module for performing dense convolution calculations according to the weight vector and the feature vector to generate a first feature; A depthwise separable convolution module for generating a second feature, where the second feature is the depthwise separable convolution weight, and for performing depthwise separable convolution calculations according to the first feature and the second feature to obtain a third feature; A storage module for storing the first feature and the third feature into the feature cache module.

2. The neural network acceleration device for pedestrian re-identification according to claim 1, characterized in that, The multi-branch network structure of the feature cache module includes a first feature cache module, two second feature cache modules, three third feature cache modules, three fourth feature cache modules and a fifth feature cache module; The first feature cache module obtains and caches the input features according to the write control signal, performs convolution calculations on the input features to generate intermediate output features and caches them in the first of the second feature cache modules. The intermediate output features generate a first convolution result, a second convolution result, a third convolution result and a fourth convolution result through Lite 3×3 convolution calculations, and stores the first convolution result into the first of the third feature cache modules, stores the second convolution result into the first of the fourth feature cache modules, stores the third convolution result into the second of the fourth feature cache modules, stores the fourth convolution result into the third of the fourth feature cache modules. The first convolution result and the second convolution result are fused to generate a first fusion result and stored in the fifth feature cache module. The first fusion result and the third convolution result are fused to generate a second fusion result and stored in the second of the third feature cache modules. The second fusion result and the fourth convolution result are fused to generate a third fusion result and stored in the second of the second feature cache modules. Convolution calculations are performed on the third fusion result to generate a fifth convolution result and stored in the third of the third feature cache modules. Residual fusion is performed on the input features and the fifth convolution result to generate the intermediate features and stored in the first feature cache module.

3. The neural network acceleration device for pedestrian re-identification according to claim 2, characterized in that, The first feature cache module and the fourth feature cache module adopt a ping-pong cache structure.

4. A neural network acceleration device for pedestrian re-identification according to claim 2, characterized in that, the storage methods of the first feature cache module, the second feature cache module, the third feature cache module, the fourth feature cache module, and the fifth feature cache module include feature width-first and feature channel-first.

5. A neural network acceleration device for pedestrian re-identification according to claim 1, characterized in that, the weight cache module adopts a ping-pong cache structure.

6. A neural network acceleration device for pedestrian re-identification according to claim 1, characterized in that, the dense convolution module includes a multiplication module and an addition tree; the multiplication module performs a multiplication operation according to the weight vector and the feature vector to obtain a product vector; the addition tree adds the product vectors to obtain the first feature.

7. A neural network acceleration device for pedestrian re-identification according to claim 1, characterized in that, it further includes a pooling module; the pooling module performs max pooling calculation or average pooling calculation according to the first feature to obtain a pooling result; the storage module stores the pooling result in the feature cache module.

8. A neural network acceleration method for pedestrian re-identification, characterized in that, the method is applied to a neural network acceleration device for pedestrian re-identification, the neural network acceleration device for pedestrian re-identification includes a control module, a cache module including a weight cache module and a feature cache module, a row cache module, a dense convolution module, a depthwise separable convolution module, and a storage module, the cache module includes a weight cache module and a feature cache module, the feature cache module is a multi-branch network structure, and the method includes: generating a read control signal and a write control signal through the control module; acquiring and caching weights through the weight cache module according to the write control signal; acquiring input features through the feature cache module according to the write control signal, and calculating and caching intermediate features according to the input features; acquiring the weights and the intermediate features through the row cache module according to the read control signal, and rearranging the weights and the intermediate features to obtain a weight vector and a feature vector; performing dense convolution calculation through the dense convolution module according to the weight vector and the feature vector to generate a first feature; generating a second feature through the depthwise separable convolution module according to the first feature, the second feature being a depthwise separable convolution weight; performing depthwise separable convolution calculation through the depthwise separable convolution module according to the first feature and the second feature to obtain a third feature; storing the first feature and the third feature in the feature cache module through the storage module.

9. A neural network acceleration method for pedestrian re-identification according to claim 8, characterized in that, the multi-branch network structure of the feature cache module includes a first feature cache module, two second feature cache modules, three third feature cache modules, three fourth feature cache modules, and a fifth feature cache module; According to the write control signal, the input features are obtained through the feature cache module, and intermediate features are calculated based on the input features and cached, including: According to the write control signal, the input features are obtained and cached through the first feature cache module; Convolution calculation is performed on the input features to generate intermediate output features and cache them in the first second feature cache module; The intermediate output features generate a first convolution result, a second convolution result, a third convolution result, and a fourth convolution result through Lite 3×3 convolution calculation, store the first convolution result in the first third feature cache module, store the second convolution result in the first fourth feature cache module, store the third convolution result in the second fourth feature cache module, and store the fourth convolution result in the third fourth feature cache module; The first convolution result and the second convolution result are fused to generate a first fusion result and stored in the fifth feature cache module; The first fusion result and the third convolution result are fused to generate a second fusion result and stored in the second third feature cache module; The second fusion result and the fourth convolution result are fused to generate a third fusion result and stored in the second second feature cache module; Convolution calculation is performed on the third fusion result to generate a fifth convolution result and stored in the third third feature cache module; Residual fusion is performed on the input features and the fifth convolution result to generate the intermediate features and store them in the first feature cache module.

10. A neural network acceleration method for pedestrian re-identification according to claim 8, wherein, the neural network acceleration device for pedestrian re-identification further includes a pooling module; the neural network acceleration method for pedestrian re-identification further includes: According to the first feature, max pooling calculation or average pooling calculation is performed through the pooling module to obtain a pooling result; The pooling result is stored in the feature cache module through the storage module.

Citation Information

Patent Citations

  • A target re-recognition method and device based on a feature selection convolution neural network

    CN109344695A

  • Efficient pedestrian re-identification method based on neural network unsupervised contrast learning

    CN111611880A