In-memory computing system and method based on product cache

Through an in-memory computing system based on product cache, the product cache and multiplexing of local features and model weights is solved, and the existing in-memory computing system is poorly scalable under high index bit width is achieved, and high-energy-efficient in-memory computing is achieved.

CN120388269APending Publication Date: 2025-07-29TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376091.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing in-memory computing systems based on lookup tables cannot improve energy efficiency by increasing the index bit width, resulting in poor scalability and cannot meet the needs of high computing efficiency, low area overhead and low computing delay.

Method used

Using an in-memory computing system based on product cache, by buffering and multiplexing the pixel values of local features with the model weight product, the content addressable memory is used to store historical product results and local feature maps, reducing the amount of multiplication and reducing the area and power consumption overhead under high index bit widths.

Benefits of technology

It effectively reduces the amount of multiplication calculation, improves data retrieval speed and energy efficiency, reduces computing power consumption, optimizes the utilization rate of computing resources, and realizes high-energy-efficient in-memory computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388269A_ABST
    Figure CN120388269A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, and provides an in-memory computing system and method based on product cache, the system comprises a cache control module, an in-memory computing module, a feature cache module and a content addressable memory, the cache control module reads a current to-be-processed local feature map from the feature cache module, the content addressable memory is searched for according to the current local feature map to be processed, a corresponding hit address is obtained based on the inquired corresponding local feature map, and the hit address is input into the in-memory calculation module; and the in-memory calculation module reads the corresponding product result cached in advance according to the hit address, and obtains a task processing result according to the product result of all the corresponding images to be subjected to task processing. According to the method, on the basis of the data locality features in the visual task, the product of the pixel values of the locality features and the model weights is cached and reused, the multiplication calculation amount is effectively reduced, and high-energy-efficiency in-memory calculation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an in-memory computing system and method based on product caching. Background Art

[0002] With the explosion of artificial intelligence applications, the parameters and complexity of AI application models are constantly increasing, and the demand for chip computing power is also getting higher and higher. At the same time, with the development of the intelligent Internet of Things, the demand for data-intensive inference tasks on a large number of low-power edge devices has surged, posing a severe challenge to how to achieve processing requirements such as high computing efficiency, low area overhead, and low computing latency on the chip of the memory computing chip.

[0003] Currently, an in-memory computing system based on a dynamic random access memory lookup table is mainly adopted. It reconfigures an electronic random access memory (eDRAM) array into a multiply-accumulate result lookup table (LUT), and allows direct encoding and refreshing operations inside the storage unit, thereby realizing high-density and high-energy-efficiency in-memory computing.

[0004] However, the throughput gain of the storage-for-computation brought by the multiply-accumulate result lookup table constructed by the above system grows linearly with its index bit width, and its area and power consumption overhead grow exponentially with the index bit width. Therefore, this method can only work under the condition of a lower index bit width, and cannot obtain better results by increasing the index bit width, and has poor scalability. Summary of the Invention

[0005] The present invention provides an in-memory computing system and method based on product caching, which is used to solve the defect that the in-memory computing system based on a lookup table in the prior art cannot further improve energy efficiency by increasing the index bit width. Based on the data locality characteristics in visual AI tasks, by caching and reusing the product of the pixel values of local features and model weights, the area and power consumption overhead under a high index bit width are effectively reduced, the multiplication calculation amount is reduced, and high-energy-efficiency in-memory computing is realized.

[0006] The present invention provides an in-memory computing system based on product caching, including a cache control module, an in-memory computing module, a feature cache module for storing a local feature map extracted from an image to be processed, and a content-addressable memory for storing the cache address of historical product results and the corresponding content of historical local feature maps, wherein: the cache control module reads the current local feature map to be processed from the feature cache module, searches the content-addressable memory according to the current local feature map to be processed, and based on the corresponding local feature map queried, obtains the corresponding hit address, and inputs the hit address into the in-memory computing module; the in-memory computing module reads the corresponding previously cached product result according to the hit address, and obtains the task processing result according to all the product results corresponding to the images to be processed.

[0007] According to an in-memory computing system based on product cache provided by the present invention, the in-memory computing module includes in-memory computing units corresponding to the weights of the corresponding layers of a preset neural network. The in-memory computing unit includes a preset cache multiplier and a preset adder tree. The cache control module includes a cache control unit, wherein: The cache control unit generates a data reading request according to the current local feature map to be processed of the image to be task-processed, and sends the data reading request to the feature cache module, and receives the local feature map returned by the feature cache module based on the data reading request. According to the input local feature map, it searches the content-addressable memory to determine whether the corresponding local feature map is found, and based on finding the corresponding local feature map, obtains the corresponding hit address, and inputs the hit address into each preset cache multiplier; The preset cache multiplier reads the corresponding product result in the product cache unit in the preset cache multiplier according to the input hit address, and inputs the read product result into the preset adder tree; The preset adder tree obtains the corresponding task processing result by using a preset accumulation algorithm according to all the input product results corresponding to the image to be task-processed.

[0008] According to an in-memory computing system based on product cache provided by the present invention, the system further includes a control module and a weight cache module for storing the weights obtained based on a preset neural network, wherein: The cache control module, based on not finding the corresponding local feature map, uses a preset cache mapping algorithm according to the unfound local feature map to obtain a first replacement address for representing the address to be written in the in-memory computing module, and encodes the unfound local feature map to generate a control signal, inputs the control signal and the first replacement address into the in-memory computing module, and sends a first notification signal to the control module; wherein, the preset cache mapping algorithm is a mapping relationship constructed in advance based on the target encoding bits of the feature map; The control module controls the weight cache module to send the weights corresponding to the unfound local feature map to the in-memory computing module according to the first notification signal; The in-memory computing module obtains the corresponding product result according to the input control signal and weights, writes the corresponding product result according to the first replacement address, and obtains the task processing result according to all the product results corresponding to the image to be task-processed.

[0009] An in-memory computing system based on product caching according to the present invention, the in-memory computing module includes in-memory computing units corresponding to the weights of the corresponding layers of a preset neural network. The in-memory computing units include a preset cache multiplier, a preset adder tree, and a 16-bit multiplier based on the Booth algorithm. The cache control module includes a cache control unit and a Booth16 encoder, where: The cache control unit, based on not querying the corresponding local feature map, according to the unqueried local feature map, uses a preset cache mapping algorithm to obtain a first replacement address, and inputs the first replacement address into each preset cache multiplier, inputs the unqueried local feature map into the Booth16 encoder, and sends a first notification signal to the control module; The Booth16 encoder, based on a preset encoding rule, encodes the unqueried local feature map to obtain a control signal for the unqueried local feature map, and inputs the control signal for the unqueried local feature map into each Booth16 multiplier; The control module, according to the first notification signal, controls the weight cache module to send the weights corresponding to the unqueried local feature map to the Booth16 multiplier; The Booth16 multiplier, according to the control signal of the unqueried local feature map and the weights corresponding to the unqueried local feature map, combines a preset weight lookup table to obtain a corresponding product result, and inputs the corresponding product result into the preset cache multiplier; where the preset weight lookup table is used to limit the adjustment range of the input weights; The preset cache multiplier, according to the first replacement address, writes the corresponding product result into the product cache unit in the preset cache multiplier, and inputs the corresponding product result into the preset adder tree; The preset adder tree, according to all the corresponding product results of the images to be processed by the task, uses a preset accumulation algorithm to obtain a corresponding task processing result.

[0010] According to an in-memory computing system based on product caching provided by the present invention, the control signals of the local feature map that are not queried include the least significant bit (LSB) and the most significant bit (MSB). The preset weight lookup table includes the weight to be updated and multiple preset adjustment ranges of the weight to be updated. The Booth16 multiplier includes: a weight update subunit that updates the preset weight lookup table based on the weight corresponding to the unqueried local feature map; a lookup table subunit that respectively looks up the updated preset weight lookup table according to the LSB and the MSB, and combines with the preset encoding rules of the previously obtained Booth16 encoder to determine the first weight adjustment range corresponding to the LSB and the second weight adjustment range corresponding to the MSB; a low-order decoder that obtains a first decoding result according to the LSB and the first weight adjustment range, in combination with a first preset shift operation and a first preset negation operation; a high-order decoder that obtains a second decoding result according to the MSB and the second weight adjustment range, in combination with a second preset shift operation and a second preset negation operation; and a comprehensive processing subunit that performs a third preset shift operation on the second decoding result, and based on the corresponding shift operation result, combines with the first decoding result to obtain a corresponding product result, and inputs the corresponding product result into a preset cache multiplier.

[0011] According to an in-memory computing system based on product caching provided by the present invention, the system further includes a partial sum buffer module, wherein: the in-memory computing module inputs the task processing result into the partial sum buffer module; and / or, the preset cache multiplier updates the content addressable memory according to the local feature map and the weight corresponding to the product result written into the product cache unit therein.

[0012] According to an in-memory computing system based on product caching provided by the present invention, the cache control module is further configured to: when reading the currently pending local feature map from the feature cache module, read the future local feature map that is separated from the currently pending local feature map by a target period; look up the content addressable memory according to the future local feature map, and end the process based on the corresponding local feature map queried.

[0013] A memory - in - computing system based on product cache according to the present invention. The system further includes a control module and a weight cache module for storing weights obtained based on a preset neural network, wherein: A cache control module, based on not querying the corresponding future local feature map, uses a preset cache mapping algorithm according to the not - queried future local feature map to obtain a second replacement address for representing the write - to address in the memory - in - computing module, encodes the not - queried future local feature map, generates a control signal for the not - queried future local feature map, inputs the control signal for the not - queried future local feature map and the second replacement address into the memory - in - computing module, and sends a second notification signal to the control module; wherein the preset cache mapping algorithm is a mapping relationship constructed in advance based on the target encoding bits of the feature map; A control module, according to the second notification signal, controls the weight cache module to send the weights corresponding to the not - queried future local feature map to the memory - in - computing module; A memory - in - computing module, according to the control signal for the not - queried future local feature map and the weights of the not - queried future local feature map, obtains the product result of the not - queried future local feature map, and writes the product result of the not - queried future local feature map according to the second replacement address.

[0014] According to an in-memory computing system based on product caching provided by the present invention, the in-memory computing module includes in-memory computing units corresponding to the weights of the corresponding layers of a preset neural network. The in-memory computing units include a preset cache multiplier, a preset adder tree, and a 16-bit multiplier based on the Booth algorithm. The cache control module includes a cache control unit and a Booth16 encoder, wherein: The cache control unit, based on the fact that the future local feature map has not been queried, according to the unqueried future local feature map, uses a preset cache mapping algorithm to obtain a second replacement address for representing the write address to be written in the in-memory computing module, and inputs the second replacement address into each preset cache multiplier, inputs the unqueried future local feature map into the Booth16 encoder, and sends a second notification signal to the control module; The Booth16 encoder, based on a preset encoding rule, encodes the unqueried future local feature map to obtain a control signal for the unqueried future local feature map, and inputs the control signal for the unqueried future local feature map into each Booth16 multiplier; The control module, according to the second notification signal, controls the weight cache module to send the weights corresponding to the unqueried future local feature map to the Booth16 multiplier; The Booth16 multiplier, according to the control signal of the unqueried future local feature map and the weights corresponding to the unqueried future local feature map, combines a preset weight lookup table to obtain a product result for the unqueried future local feature map, and inputs the corresponding product result into the preset cache multiplier; wherein, the preset weight lookup table is used to limit the adjustment range of the input weights; The preset cache multiplier, according to the second replacement address, writes the product result of the unqueried future local feature map into the product cache unit in the preset cache multiplier.

[0015] The present invention also provides an in-memory computing method based on product caching, which is applied to the in-memory computing system based on product caching as described above. The method includes: reading the currently to-be-processed local feature map, where the local feature map is extracted in advance based on the to-be-task-processed image; according to the currently to-be-processed local feature map, searching the content-addressable memory, and based on the corresponding local feature map queried, obtaining the corresponding hit address; wherein, the content-addressable memory is used to store the cache addresses of historical product results and the corresponding historical local feature maps; according to the hit address, reading the corresponding product result cached in advance, and obtaining the task processing result according to all the product results corresponding to the to-be-task-processed images.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the in-memory computing method based on product caching as described above is implemented.

[0017] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the in-memory computing method based on product caching as described in any one of the above.

[0018] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the in-memory computing method based on product caching as described in any one of the above.

[0019] The in-memory computing system and method based on product caching provided by the present invention store local feature maps to be processed through a feature caching module, so as to be quickly accessed by a cache control module, reduce unnecessary data access, improve energy efficiency, reduce memory access latency, and the cache control module searches a content-addressable memory according to the read local feature map to directly access relevant data by matching the content, reduce the area and power consumption overhead under a high index bit width, effectively improve the speed and energy efficiency of data retrieval, and based on being able to find the corresponding product result, obtain a hit address, so that the in-memory computing unit reads the corresponding product result according to the hit address, avoid repeated calculation of partial products, reduce the multiplication calculation amount, reduce the data calculation time, speed up the processing speed, and reduce the computing power consumption; in addition, based on the data locality characteristics in vision AI tasks, cache the pixel values of the local feature map multiplied by the model weights and reuse them in the foregoing manner to reduce repeated calculations and optimize the utilization rate of computing resources, and achieve high-energy-efficiency in-memory computing. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 is one of the structural schematic diagrams of the in-memory computing system based on product caching provided by the present invention; Figure 2 is another structural schematic diagram of the in-memory computing system based on product caching provided by the present invention; Figure 3 is the control flow chart of the cache control module provided by the present invention; Figure 4 is the control flow chart of the Booth16 multiplier provided by the present invention; Figure 5 is the flow schematic diagram of the in-memory computing method based on product caching provided by the present invention; Figure 6 is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners

[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] Figure 1 is a schematic flowchart of an in-memory computing system based on product caching provided by the present invention. As Figure 1 shown, the system includes a cache control module, an in-memory computing module, a feature cache module for storing local feature maps extracted from images to be processed, and a content-addressable memory for storing the cache addresses of historical product results and the corresponding historical local feature maps, where: The cache control module reads the currently to-be-processed local feature map from the feature cache module, searches the content-addressable memory according to the currently to-be-processed local feature map, and based on the corresponding local feature map found, obtains the corresponding hit address, and inputs the hit address into the in-memory computing module; The in-memory computing module reads the corresponding previously cached product result according to the hit address, and obtains the task processing result according to all the product results of the corresponding images to be processed.

[0024] It should be noted that by using a content-addressable memory (Content-Addressable Memory, abbreviated as CAM) to store the addresses of historical product results and the corresponding historical local feature maps, it is convenient to directly perform address search according to the local feature maps subsequently, thereby improving the efficiency of data retrieval.

[0025] Specifically, referring to Figures 2 - 3, the in-memory computing module includes in-memory computing units corresponding to the weights of the corresponding layers of the preset neural network. The in-memory computing units include a preset cache multiplier and a preset adder tree. The cache control module includes a cache control unit, where: The cache control unit generates a data reading request according to the currently to-be-processed local feature map of the image to be task-processed, and sends the data reading request to the feature cache module, and receives the local feature map returned by the feature cache module based on the data reading request. According to the input local feature map, it searches the content-addressable memory to determine whether the corresponding local feature map is found, and based on finding the corresponding local feature map, obtains the corresponding hit address, and inputs the hit address into each preset cache multiplier; The preset cache multiplier reads the corresponding product result in the product cache unit in the preset cache multiplier according to the input hit address, and inputs the read product result into the preset adder tree; The preset adder tree uses the preset accumulation algorithm to obtain the corresponding task processing result according to all the input product results corresponding to the image to be task-processed.

[0026] It should be added that the preset cache control module can be selected according to actual design requirements, such as PRE 3 cache controller, etc. Assuming that the basic operator of this system is a vector dot product with a length of 32, corresponding to 32 product caches. During the execution of the calculation, a miss in any one of the product caches will cause the entire calculation circuit to block. Through PRE 3 cache controller, the occurrence of blocking can be effectively reduced and the calculation throughput can be improved. When the basic operator of the system is a vector dot product with a length of 32, the input channels of the preset cache control unit are 32 and the output channels are 16. Correspondingly, the preset cache control unit searches the content-addressable memory based on the local feature maps obtained from each input channel to determine whether each local feature map is pre-cached, and based on being able to find the corresponding local feature map in the content-addressable memory, determines that the corresponding local feature map has been pre-cached, determines the hit address of each local feature map, and inputs each hit address into the corresponding in-memory computing unit through its output channels.

[0027] In addition, the cache control unit needs to read the local feature map of the image to be task-processed cached in the feature cache module based on a preset reading frequency, so as to process the next local feature map in a timely manner after processing the current local feature map, thereby realizing the processing of the image to be task-processed. The preset reading frequency can be set according to actual design requirements such as actual computing power resources and the duration of processing a single read local feature map, and no further limitation is made here.

[0028] In addition, each in-memory computing unit needs to process the cumulative multiplication and operation of the corresponding input currently to-be-processed local feature map and the corresponding weights, that is, the vector dot product , where A[i] represents the current local feature map to be processed each time an input is received, and W[i] represents the weights allocated to the corresponding in-memory computing unit. The cache control unit reads A[i] from the feature cache module multiple times and searches the content-addressable memory. When it is determined that the corresponding local feature map exists in the content-addressable memory, the preset cache multiplier directly reads the corresponding product result according to the hit address. , and sends to the preset adder tree. The preset adder tree accumulates according to all , i = 0,,, 31, to obtain the cumulative multiplication and operation result, and uses it as the task processing result of the corresponding layer of the preset neural network corresponding to the corresponding in-memory computing unit.

[0029] In an alternative embodiment, the CAM further includes the cumulative multiplication and result output by each in-memory computing unit. Therefore, the system further includes a partial sum buffer module, where: the in-memory computing module inputs the task processing result into the partial sum buffer module.

[0030] In an alternative embodiment, the currently obtained local feature map may be a feature that has not appeared before or a feature for which the corresponding product result has not been cached before. The cache control module cannot find the corresponding local feature map from the content-addressable memory based on the currently to-be-processed local feature map. Therefore, the system further includes a control module and a weight cache module for storing the weights obtained based on the preset neural network. Specifically: the cache control module, based on the failure to query the corresponding local feature map, uses the preset cache mapping algorithm according to the unqueried local feature map to obtain a first replacement address for characterizing the write address in the in-memory computing module, encodes the unqueried local feature map to generate a control signal, inputs the control signal and the first replacement address into the in-memory computing module, and sends a first notification signal to the control module; where the preset cache mapping algorithm is a mapping relationship constructed based on the target encoding bits of the feature map before; the control module, according to the first notification signal, controls the weight cache module to send the weights corresponding to the unqueried local feature map to the in-memory computing module; the in-memory computing module, according to the input control signal and weights, obtains the corresponding product result, writes the corresponding product result according to the first replacement address, and obtains the task processing result according to all the product results corresponding to the images to be task-processed.

[0031] Specifically, the in-memory computing module includes in-memory computing units corresponding to the weights of the corresponding layers of the preset neural network. The in-memory computing units include a preset cache multiplier, a preset adder tree, and a 16-bit multiplier based on the Booth algorithm. The cache control module includes a cache control unit and a Booth16 encoder, where: The cache control unit, based on the non-detection of the corresponding local feature map, uses the preset cache mapping algorithm to obtain the first replacement address according to the non-detected local feature map, inputs the first replacement address into each preset cache multiplier, inputs the non-detected local feature map into the Booth16 encoder, and sends a first notification signal to the control module; The Booth16 encoder encodes the non-detected local feature map based on the preset encoding rule to obtain the control signal of the non-detected local feature map, and inputs the control signal of the non-detected local feature map into each Booth16 multiplier; The control module controls the weight cache module to send the weights corresponding to the non-detected local feature map to the Booth16 multiplier according to the first notification signal; The Booth16 multiplier obtains the corresponding product result by combining the control signal of the non-detected local feature map and the weights corresponding to the non-detected local feature map with the preset weight lookup table, and inputs the corresponding product result into the preset cache multiplier; where the preset weight lookup table is used to limit the adjustment range of the input weights; The preset cache multiplier writes the corresponding product result into the product cache unit in the preset cache multiplier according to the first replacement address, and inputs the corresponding product result into the preset adder tree; The preset adder tree obtains the corresponding task processing result by using the preset accumulation algorithm according to all the product results of the corresponding images to be task-processed input.

[0032] It should be noted that based on the good data locality of the feature maps in the vision AI task, the product results of the local feature maps of the images to be task-processed and the corresponding weights can be cached in the product cache unit of the preset cache multiplier, and the local feature maps and the corresponding weights corresponding to the cached product results can be updated to the content-addressable memory, so that when processing the local feature maps of the corresponding images to be task-processed, it is possible to determine whether the corresponding product result is cached based on looking up the content-addressable memory, thereby facilitating directly reading the corresponding product result from the preset cache multiplier. Through the data locality feature in the vision AI task, the product results of the cached hot-spot feature map data and the corresponding weight data are cached, avoiding repeated calculation of this part of the product, realizing the reduction of computing power consumption, and by calculating the corresponding product result in the above manner when the corresponding local feature map cannot be queried from the content-addressable memory, and updating the content-addressable memory in a timely manner, so as to achieve a high cache hit rate with a small-capacity product cache, and further achieve the effect of saving power.

[0033] It should be added that the cache control unit further includes a cache replacement address generator. Based on the failure to query the corresponding local feature map, the cache replacement address generator is used to generate the first replacement address of the unqueried local feature map. A preset cache mapping algorithm is configured in the cache replacement address generator. Further, the preset cache mapping algorithm can be configured according to actual design requirements. For example, if the actual local feature map is 8-bit encoded bit data, its lower 4 bits can be used as the corresponding replacement address, so as to determine the write position of the product cache unit in the corresponding preset cache multiplier according to this replacement address, such as which row of the corresponding product cache unit to write to. No further limitation is made here.

[0034] In addition, the preset encoding rule can be set according to the actually adopted Booth16 encoder and actual encoding requirements. For example, if the local feature map is 8-bit data, then based on the preset encoding rule, the unqueried local feature map is encoded, including: according to the unqueried local feature map, padding 0 to the lower four bits and the last bit to obtain the least significant bit (LSB), and obtaining the most significant bit (MSB) according to the higher 5 bits. The specific encoding can be determined according to the actually designed encoding rule. No further limitation is made here.

[0035] Correspondingly, after calculating the product result of the unqueried local feature map by the above method, it is also necessary to update the corresponding content addressable memory according to the product result of the unqueried local feature map. Therefore, after obtaining the above task processing result, the preset cache multiplier updates the content addressable memory according to the local feature map and weight corresponding to the product result written into the product cache unit therein. Thus, based on the good data locality of the feature map of the vision AI task, the product results of each local feature map and the corresponding weight of the image to be processed by the task are cached in the product cache unit of the preset cache multiplier, and the local feature map and the corresponding weight corresponding to the cached product result are updated to the content addressable memory, so that when processing the local feature map of the corresponding image to be processed by the task subsequently, it is possible to determine whether the corresponding product result is cached based on looking up the content addressable memory, and then it is convenient to directly read the corresponding product result from the preset cache multiplier to achieve a high cache hit rate through a small-capacity product cache and achieve the effect of power saving. In addition, the in-memory computing module inputs the task processing result into the partial sum buffer module.

[0036] In an alternative embodiment, in order to reduce the computational overhead when there is a cache miss, the control signal of the unqueried local feature map includes the least significant bit (LSB) and the most significant bit (MSB). The preset weight lookup table includes the weight to be updated and multiple preset adjustment ranges of the weight to be updated. The Booth16 multiplier includes: a weight update subunit that updates the preset weight lookup table based on the weight corresponding to the unqueried local feature map; a table lookup subunit that, according to the LSB and the MSB respectively, combines the preset encoding rules of the previously obtained Booth16 encoder to look up the updated preset weight lookup table and determines the first weight adjustment range corresponding to the LSB and the second weight adjustment range corresponding to the MSB; a low-order decoder that, according to the LSB and the first weight adjustment range, combines the first preset shift operation and the first preset negation operation to obtain a first decoding result; a high-order decoder that, according to the MSB and the second weight adjustment range, combines the second preset shift operation and the second preset negation operation to obtain a second decoding result; a comprehensive processing subunit that performs a third preset shift operation on the second decoding result and, based on the corresponding shift operation result, combines the first decoding result to obtain a corresponding product result and inputs the corresponding product result into the preset cache multiplier.

[0037] It should be added that the adjustment range of the preset weight lookup table can be set according to actual design requirements. For example, referring to Figure 4 , the adjustment range can include 1W, 3W, 5W, and 7W, where W represents the weight. By using the weight corresponding to the unqueried local feature map to update the weight in the preset weight lookup table, the preset weight lookup table is updated. In addition, the preset encoding rules are also used to define the adjustment range corresponding to specific bit values of the LSB and the MSB. For example, when it is 11, the corresponding adjustment range is 1W; when it is 00, the corresponding adjustment range is 3W; when it is 01, the corresponding adjustment range is 5W; when it is 10, the corresponding adjustment range is 7W. It can be specifically set according to actual design requirements and will not be further limited here. In addition, the Booth16 multiplier is implemented by adding a 1 / 3 / 5 / 7 lookup table and using a 1-write-port 2-read-port 6T-SRAM array. Therefore, only 1 adder is required to complete the multiplication operation in a single cycle, achieving lower power consumption and a smaller area.

[0038] Furthermore, continuing to refer to Figure 4, the Booth16 multiplier further includes a ROM-based read 0 circuit subunit. The ROM-based read 0 circuit subunit includes a word line low level WLL, a word line read signal WLR, a bit line BL, and a complementary bit line BLB. Among them: WLL selects an address, WLR reads 0 data from the corresponding position in the ROM according to the address selected by WLL, and transmits it through BL. BLB is used to transmit differential signals when BL transmits data to improve the anti-interference ability, so as to read the 0 bit, which serves as an external zero-weight bit. Without changing the value of the multiplier, it helps the Booth16 algorithm process the highest bit of the multiplier more accurately.

[0039] In addition, the first preset shift operation and the second preset shift operation can select the corresponding shift times for left shift operations from 0 to 3 times according to the corresponding high-order decoder and low-order decoder and their respective actual decoding requirements. The third preset shift operation can be configured according to the result output by the low-order decoder and the actual design requirements, such as a left shift operation of shifting 4 times to the left, etc. No further details are provided here. Finally, the comprehensive processing subunit obtains the corresponding product result by adding the result of the third preset shift operation to the first decoding result.

[0040] While the cache control module reads the currently pending local feature map, it can also read the local feature maps within a certain period in the future of the currently pending local feature map to perform pre-lookup on the future local feature maps, so as to facilitate determining whether the product results corresponding to the future local feature maps have been pre-cached, and in the case of non-caching, perform preprocessing to cache the corresponding product results, only to facilitate directly reading the corresponding cached product results based on the lookup when the corresponding future local feature map is used as the currently pending local feature map in the future, so as to effectively reduce the occurrence of blocking and improve the computing throughput.

[0041] Correspondingly, the cache control module is further configured to: when reading the currently pending local feature map from the feature cache module, read the future local feature map that is separated from the currently pending local feature map by a target period; search the content addressable memory according to the future local feature map, and end the process based on the queried corresponding local feature map.

[0042] It should be added that the target period can be configured according to the data stream in the actual neural network, the actual computing power of the system, and the processing capacity. The data stream in the neural network is fixed and predictable. Therefore, the target period can be the next two periods of the current local feature map to be processed, corresponding to the prefetch queue positions Act[N], Act[N+1], and Act[N+2]. The prefetch queue can be a FIFO (First In First Out) queue. Act[N] is used as the current local feature map to be processed, and Act[N+2] is used as the future local feature map for pre-lookup and preprocessing. If the cache control module detects a miss for two periods, it will use the above-mentioned Booth16 encoder and Booth16 multiplier for pre-update operations, thereby achieving the effect of reducing stalls and increasing the computing throughput rate. The specific implementation method can be referred to the foregoing description and will not be repeated here.

[0043] Specifically, the system further includes a control module and a weight cache module for storing weights obtained based on a preset neural network, where: The cache control module, based on the non-detection of the corresponding future local feature map, uses a preset cache mapping algorithm according to the non-detected future local feature map to obtain a second replacement address for representing the write address in the in-memory computing module, and encodes the non-detected future local feature map to generate a control signal for the non-detected future local feature map. The control signal and the second replacement address of the non-detected future local feature map are input into the in-memory computing module, and a second notification signal is sent to the control module; where the preset cache mapping algorithm is a mapping relationship constructed in advance based on the target encoding bits of the feature map; The control module, according to the second notification signal, controls the weight cache module to send the weights corresponding to the non-detected future local feature map to the in-memory computing module; The in-memory computing module, according to the control signal of the non-detected future local feature map and the weights of the non-detected future local feature map, obtains the product result of the non-detected future local feature map and writes the product result of the non-detected future local feature map according to the second replacement address. The specific processing method can be referred to the foregoing description and will not be further described here.

[0044] It should be noted that after calculating the product result of the non-detected future local feature map using the above method, in order to facilitate directly finding it based on the cache when the future local feature map is used as the current local feature map to be processed later, it is also necessary to update the corresponding content addressable memory in a timely manner according to the product result of the non-detected future local feature map. This will not be repeated here, and since it is preprocessing of the non-detected future local feature map, it is not necessary to input the product result of the non-detected future local feature map into the corresponding preset adder tree for accumulation operation.

[0045] In addition, the in-memory computing module includes in-memory computing units corresponding to the weights of the corresponding layers of the preset neural network. The in-memory computing unit includes a preset cache multiplier, a preset adder tree, and a 16-bit multiplier based on the Booth algorithm. The cache control module includes a cache control unit and a Booth16 encoder, where: The cache control unit, based on the fact that the future local feature map has not been queried, uses the preset cache mapping algorithm according to the unqueried future local feature map to obtain a second replacement address for representing the write address in the in-memory computing module, and inputs the second replacement address into each preset cache multiplier, inputs the unqueried future local feature map into the Booth16 encoder, and sends a second notification signal to the control module; The Booth16 encoder encodes the unqueried future local feature map based on the preset encoding rule to obtain the control signal of the unqueried future local feature map, and inputs the control signal of the unqueried future local feature map into each Booth16 multiplier; The control module, according to the second notification signal, controls the weight cache module to send the weights corresponding to the unqueried future local feature map to the Booth16 multiplier; The Booth16 multiplier, according to the control signal of the unqueried future local feature map and the weights corresponding to the unqueried future local feature map, combines the preset weight lookup table to obtain the product result of the unqueried future local feature map, and inputs the corresponding product result into the preset cache multiplier; where the preset weight lookup table is used to limit the adjustment range of the input weights; The preset cache multiplier writes the product result of the unqueried future local feature map into the product cache unit in the preset cache multiplier according to the second replacement address. The specific processing method can refer to the foregoing description and will not be further described here.

[0046] In an alternative embodiment, the system includes a weight cache module, a feature cache module, 1 cache control module, and 16 in-memory computing units. The cache control unit includes 32 input channels ICHs and 16 output channels. The cache control module includes a cache control unit and a Booth16 encoder. The weight matrix is stored in the weight cache module , and the feature cache module inputs the feature map vector into the cache control unit , and performs matrix-vector multiplication operation to obtain the output .

[0047] Furthermore, each in-memory computing unit includes a preset cache multiplier, a preset adder tree, and a Booth16 multiplier, and is responsible for performing a vector dot product with a length of 32 , the preset cache multiplier is responsible for completing 32 multiplication operations and obtaining 32 16-bit product results; the preset adder tree is responsible for adding up the 32 16-bit product results to obtain a 24-bit multiply-accumulate result. Each preset cache multiplier contains 32 product cache units and 8 Booth16 multipliers, and every 4 product cache units share 1 Booth16 multiplier.

[0048] In addition, the cache control module includes 32 product cache controllers and 8 Booth16 encoders. Each product cache controller contains a prefetch FIFO with a depth of 3, a CAM with a dual-query port, and a cache replacement address generator.

[0049] Furthermore, based on the above method, the data stream when the cache hits is obtained, specifically: the cache control module reads the feature map data ACT[N + 2] required for calculation 2 cycles after reading from the feature cache module; the cache control module queries the local feature map ACT[N] required for calculation in this cycle in the CAM; the cache control module sends the hit address of the cache hit to the product cache unit of the preset cache multiplier corresponding to the input channel in 16 in-memory computing units; the product cache unit reads out the previously cached product result according to the hit address; after the caches of all 32 input channels hit, each computing core sends the product results of 32 input channels to the preset adder tree, and the preset adder tree completes the accumulation operation and inputs the result into the partial sum buffer module.

[0050] In addition, based on the above method, the data stream when the cache misses is obtained, specifically: the cache control module reads the feature map data ACT[N + 2] required for calculation 2 cycles after reading from the feature cache module; the cache control module does not query ACT[N + 2] in the CAM; the cache replacement address generator determines the cache address to be replaced according to a specific replacement policy and sends it to the product cache unit of the preset cache multiplier corresponding to the input channel in 16 in-memory computing units; the Booth16 encoder generates a control signal according to the value of ACT[N + 2] according to the encoding rules of the Booth16 algorithm and sends it to the Booth16 multiplier corresponding to the input channel in 16 in-memory computing units; the Booth16 multiplier generates a product result according to the Booth16 control signal; the product cache unit writes the product result generated by the Booth16 multiplier into the replacement address provided by the cache control module.

[0051] In summary, in the embodiment of the present invention, the feature cache module stores the local feature map to be processed, so as to be quickly accessed by the cache control module, reduce unnecessary data access, improve energy efficiency, reduce memory access latency, and the cache control module searches the content addressable memory according to the read local feature map to directly access relevant data by matching the content, reduce the area and power consumption overhead under high index bit widths, effectively improve the speed and energy efficiency of data retrieval, and based on being able to find the corresponding product result, obtain the hit address, so that the in-memory computing unit reads the corresponding product result according to the hit address, avoid repeated calculation of partial products, reduce the multiplication calculation amount, reduce the data calculation time, speed up the processing speed, and reduce the calculation power consumption; in addition, based on the data locality feature in the vision AI task, the pixel values of the local feature map and the product of the model weights are cached and reused in the foregoing manner to reduce repeated calculation and optimize the utilization rate of computing resources, realizing high-energy-efficiency in-memory computing.

[0052] The in-memory computing method based on product cache provided by the present invention will be described below. The in-memory computing method based on product cache described below can be correspondingly referred to the in-memory computing system based on product cache described above.

[0053] Figure 5 The flowchart of an in-memory computing method based on product cache is shown, which is applied to any of the in-memory computing systems based on product cache as above. The method includes: S51, read the current local feature map to be processed, where the local feature map is extracted in advance based on the image to be task-processed; S52, search the content addressable memory according to the current local feature map to be processed, and obtain the corresponding hit address based on the queried corresponding local feature map; wherein, the content addressable memory is used to store the cache addresses of historical product results and the corresponding historical local feature maps; S53, read the corresponding product result cached in advance according to the hit address, and obtain the task processing result according to all the product results corresponding to the images to be task-processed.

[0054] It should be noted that the specific principle of the embodiment of the present invention is the same as the principle of the above system embodiment. For details, please refer to the foregoing system embodiment, and no more detailed explanations will be given here.

[0055] Figure 6 The schematic physical structure diagram of an electronic device is illustrated, as Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute an in-memory computing method based on product caching. The method includes: reading a currently to-be-processed local feature map, where the local feature map is previously extracted based on an image to be task-processed; according to the currently to-be-processed local feature map, searching a content-addressable memory, and obtaining a corresponding hit address based on the corresponding local feature map found; where the content-addressable memory is used to store the cache addresses of historical product results and the corresponding historical local feature maps; according to the hit address, reading the corresponding previously cached product result, and obtaining a task processing result according to all the product results corresponding to the images to be task-processed.

[0056] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0057] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the in-memory computing method based on product caching provided by the above-mentioned various methods. The method includes: reading a currently to-be-processed local feature map, where the local feature map is previously extracted based on an image to be task-processed; according to the currently to-be-processed local feature map, searching a content-addressable memory, and obtaining a corresponding hit address based on the corresponding local feature map found; where the content-addressable memory is used to store the cache addresses of historical product results and the corresponding historical local feature maps; according to the hit address, reading the corresponding previously cached product result, and obtaining a task processing result according to all the product results corresponding to the images to be task-processed.

[0058] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements an in-memory computing method based on product caching provided by the above-mentioned various methods. The method includes: reading a currently to-be-processed local feature map, where the local feature map is extracted in advance based on an image to be task-processed; according to the currently to-be-processed local feature map, searching a content-addressable memory, and based on the corresponding local feature map found, obtaining a corresponding hit address; wherein the content-addressable memory is used to store the cache addresses of historical product results and the corresponding historical local feature maps; according to the hit address, reading the corresponding product result cached in advance, and obtaining a task processing result based on all the product results corresponding to the images to be task-processed.

[0059] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0060] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0061] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An in-memory computing system based on product caching, characterized in that It includes a cache control module, an in-memory computing module, a feature cache module for storing local feature maps extracted from images to be task-processed, and a content-addressable memory for storing cache addresses of historical product results and contents of corresponding historical local feature maps, where: The cache control module reads the currently to-be-processed local feature map from the feature cache module, searches the content-addressable memory according to the currently to-be-processed local feature map, obtains a corresponding hit address based on the queried corresponding local feature map, and inputs the hit address into the in-memory computing module; The in-memory computing module reads the corresponding pre-cached product result according to the hit address, and obtains a task processing result according to all product results corresponding to the images to be task-processed.

2. The in-memory computing system based on product caching according to claim 1, wherein The in-memory computing module includes in-memory computing units corresponding to weights of corresponding layers of a preset neural network. The in-memory computing unit includes a preset cache multiplier and a preset adder tree. The cache control module includes a cache control unit, where: The cache control unit generates a data read request according to the currently to-be-processed local feature map of the image to be task-processed, sends the data read request to the feature cache module, receives the local feature map returned by the feature cache module based on the data read request, searches the content-addressable memory according to the input local feature map, determines whether a corresponding local feature map is queried, obtains a corresponding hit address based on the queried corresponding local feature map, and inputs the hit address into each of the preset cache multipliers; The preset cache multiplier reads the corresponding product result in the product cache unit in the preset cache multiplier according to the input hit address, and inputs the read product result into the preset adder tree; The preset adder tree obtains a corresponding task processing result according to all product results corresponding to the images to be task-processed by using a preset accumulation algorithm.

3. The in-memory computing system based on product cache according to claim 1, characterized in that, The system further includes a control module and a weight cache module for storing weights obtained based on a preset neural network, where: Based on the non-query of a corresponding local feature map, the cache control module obtains a first replacement address for characterizing the to-be-written address in the in-memory computing module according to the non-query local feature map by using a preset cache mapping algorithm, encodes the non-query local feature map to generate a control signal, inputs the control signal and the first replacement address into the in-memory computing module, and sends a first notification signal to the control module; where the preset cache mapping algorithm is a mapping relationship constructed in advance based on target encoding bits of the feature map; The control module controls the weight cache module to send the weight corresponding to the non-query local feature map to the in-memory computing module according to the first notification signal; The in-memory computing module obtains a corresponding product result according to the input control signal and weight, writes the corresponding product result according to the first replacement address, and obtains a task processing result according to all product results corresponding to the images to be task-processed.

4. The in-memory computing system based on product cache according to claim 3, characterized in that: The in-memory computing module includes in-memory computing units corresponding to the weights of the corresponding layers of a preset neural network. The in-memory computing units include a preset cache multiplier, a preset adder tree, and a 16-bit multiplier based on the Booth algorithm. The cache control module includes a cache control unit and a Booth16 encoder, where: The cache control unit, based on the non-detection of the corresponding local feature map, uses a preset cache mapping algorithm to obtain a first replacement address according to the non-detected local feature map, inputs the first replacement address into each of the preset cache multipliers, inputs the non-detected local feature map into the Booth16 encoder, and sends a first notification signal to the control module; The Booth16 encoder encodes the non-detected local feature map based on a preset encoding rule to obtain a control signal for the non-detected local feature map, and inputs the control signal for the non-detected local feature map into each of the Booth16 multipliers; The control module, according to the first notification signal, controls the weight cache module to send the weights corresponding to the non-detected local feature map to the Booth16 multiplier; The Booth16 multiplier, according to the control signal for the non-detected local feature map and the weights corresponding to the non-detected local feature map, combines a preset weight lookup table to obtain a corresponding product result, and inputs the corresponding product result into the preset cache multiplier; wherein, the preset weight lookup table is used to define the adjustment range of the input weights; The preset cache multiplier writes the corresponding product result into the product cache unit in the preset cache multiplier according to the first replacement address, and inputs the corresponding product result into the preset adder tree; The preset adder tree, according to all the corresponding product results for the image to be task-processed, uses a preset accumulation algorithm to obtain a corresponding task processing result.

5. The in-memory computing system based on product cache according to claim 4, characterized in that: The control signal for the non-detected local feature map includes the least significant bit (LSB) and the most significant bit (MSB). The preset weight lookup table includes the weights to be updated and multiple preset adjustment ranges for the weights to be updated. The Booth16 multiplier includes: A weight update sub-unit that updates the preset weight lookup table based on the weights corresponding to the non-detected local feature map; A table lookup sub-unit that respectively looks up the updated preset weight lookup table according to the LSB and the MSB, in combination with the preset encoding rule of the previously obtained Booth16 encoder, to determine a first weight adjustment range corresponding to the LSB and a second weight adjustment range corresponding to the MSB; A low-order decoder that, according to the LSB and the first weight adjustment range, in combination with a first preset shift operation and a first preset inversion operation, obtains a first decoding result; A high-order decoder that, according to the MSB and the second weight adjustment range, in combination with a second preset shift operation and a second preset inversion operation, obtains a second decoding result; The comprehensive processing sub-unit performs a third preset shift operation on the second decoding result, combines the first decoding result based on the corresponding shift operation result to obtain a corresponding product result, and inputs the corresponding product result into the preset cache multiplier.

6. The in-memory computing system based on product cache according to claim 4, characterized in that: The system further includes a partial sum buffer module, where: The in-memory computing module inputs the task processing result into the partial sum buffer module; and / or, The preset cache multiplier updates the content-addressable memory according to the local feature map and weight corresponding to the product result written into the product cache unit therein.

7. The in-memory computing system based on product cache according to claim 1, wherein The cache control module is further configured to: When reading the currently to-be-processed local feature map from the feature cache module, read the future local feature map that is separated from the currently to-be-processed local feature map by a target period; According to the future local feature map, search the content-addressable memory, and end the process based on the corresponding local feature map found.

8. The in-memory computing system based on product cache according to claim 7, characterized in that: The system further includes a control module and a weight cache module for storing weights obtained based on a preset neural network, where: Based on not finding the corresponding future local feature map, the cache control module uses a preset cache mapping algorithm according to the not-found future local feature map to obtain a second replacement address for representing the write address in the in-memory computing module, encodes the not-found future local feature map to generate a control signal for the not-found future local feature map, inputs the control signal for the not-found future local feature map and the second replacement address into the in-memory computing module, and sends a second notification signal to the control module; wherein, the preset cache mapping algorithm is a mapping relationship previously constructed based on the target encoding bits of the feature map; The control module controls the weight cache module to send the weight corresponding to the not-found future local feature map to the in-memory computing module according to the second notification signal; The in-memory computing module obtains the product result of the not-found future local feature map according to the control signal for the not-found future local feature map and the weight of the not-found future local feature map, and writes the product result of the not-found future local feature map at the second replacement address.

9. The in-memory computing system based on product cache according to claim 8, wherein The in-memory computing module includes in-memory computing units corresponding to the weights of the corresponding layers of the preset neural network. The in-memory computing unit includes a preset cache multiplier, a preset adder tree, and a 16-bit multiplier based on the Booth algorithm. The cache control module includes a cache control unit and a Booth16 encoder, where: Based on not finding the future local feature map, the cache control unit uses a preset cache mapping algorithm according to the not-found future local feature map to obtain a second replacement address for representing the write address in the in-memory computing module, inputs the second replacement address into each of the preset cache multipliers, inputs the not-found future local feature map into the Booth16 encoder, and sends a second notification signal to the control module; The Booth16 encoder encodes the unqueried future local feature map based on a preset encoding rule to obtain a control signal for the unqueried future local feature map, and inputs the control signal for the unqueried future local feature map into each of the Booth16 multipliers; The control module controls the weight cache module to send the weights corresponding to the unqueried future local feature map to the Booth16 multiplier according to the second notification signal; The Booth16 multiplier obtains a product result for the unqueried future local feature map according to the control signal for the unqueried future local feature map and the weights corresponding to the unqueried future local feature map, in combination with a preset weight lookup table, and inputs the corresponding product result into the preset cache multiplier; wherein, the preset weight lookup table is used to define the adjustment range of the input weights; The preset cache multiplier writes the product result for the unqueried future local feature map into a product cache unit in the preset cache multiplier according to the second replacement address.

10. A method for in-memory computing based on product caching, which is applied to the in-memory computing system based on product caching as described in any one of claims 1-9, and is characterized in that, The method includes: Reading a currently to-be-processed local feature map, which is previously extracted based on an image to be task-processed; Searching a content addressable memory according to the currently to-be-processed local feature map, and obtaining a corresponding hit address based on the queried corresponding local feature map; wherein, the content addressable memory is used to store cache addresses of historical product results and corresponding historical local feature maps; Reading a corresponding previously cached product result according to the hit address, and obtaining a task processing result according to all product results corresponding to the images to be task-processed.