Computing system
By introducing a computing unit of a storage device into the computing system, distributing computing tasks and completing calculations in conjunction with the processing equipment, the problems of large data handling volume and high power consumption are solved, and the computing speed is improved.
Patent Information
- Application Number
- CN202111051844.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-09-08
AI Technical Summary
There are problems in existing computing systems where data handling volume is large, power consumption is high, and calculation speed is difficult to improve.
By introducing a computing unit of a storage device into the computing system, computing tasks are distributed and calculations are completed in concert between the storage device and the processing device, data handling and power consumption are reduced.
It realizes reducing data handling, reducing power consumption and improving processing speed.
Smart Images

Figure CN113849454B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of data processing, and more particularly, to a computing system. Background Art
[0002] This section aims to provide background or context for the embodiments of the present disclosure recited in the claims. The description herein is not admitted to be prior art merely by virtue of its inclusion in this section.
[0003] In existing computing systems, the entire computing process is usually completed in a processing chip. Taking an AI computing system as an example, all the original data usually needs to be read from a non-volatile memory, such as a solid state disk (SSD) or a flash memory, into a random access memory, such as a double data rate (DDR) synchronous dynamic random access memory, and then the original data is read into an AI chip for AI computing processing by the AI chip.
[0004] However, this method has a large amount of data transfer and a large amount of computation in the processing chip, resulting in high power consumption and difficulty in improving the computing speed. Summary of the Invention
[0005] In view of this, at least one computing system is provided in the embodiments of the present disclosure to reduce transfer, lower power consumption, and improve processing speed during data computing and processing.
[0006] The present disclosure provides a computing system, including a first processing device, a second processing device, and a storage device. The storage device includes a computing unit. The first processing device is configured to distribute computing tasks to the storage device and the second processing device. The storage device is configured to use the computing unit to perform computing processing on the stored data in the storage device according to the received computing tasks to obtain intermediate processing results. The second processing device is configured to receive the data to be processed corresponding to the intermediate processing results sent by the storage device and perform computing processing on the data to be processed according to the received computing tasks to obtain target processing results, where the data to be processed includes a part of the stored data.
[0007] In some embodiments, the storage device includes a first memory and a second memory. Among them, the first memory includes a first computing unit; the second memory includes a second computing unit; the first memory is configured to use the first computing unit to perform computational processing on the stored data in the first memory to obtain a first processing result; the second memory is configured to receive and store the first data to be processed corresponding to the first processing result sent by the first memory, where the first data to be processed includes a part of the stored data; and is configured to use the second computing unit to perform computational processing on the first data to be processed in the second memory to obtain the intermediate processing result.
[0008] In some embodiments, index information of the stored data in the first memory is stored in the second memory; when the computing task includes a retrieval task, the second memory is configured to use the second computing unit to determine target index information matching the retrieved data in the index information; the first processing device is configured to obtain the data to be processed corresponding thereto in the first memory according to the target index information, and send the data to be processed to the second processing device.
[0009] In some embodiments, the index information of the stored data in the second memory includes first feature information of the stored data and storage information of the stored data in the first memory, and the first feature information of the stored data is associated with the storage information; the second memory is configured to use the second computing unit to compare second feature information of the retrieved data with the first feature information of the stored data in the second memory, and determine target feature information according to the comparison result; the first processing device is configured to obtain the data to be processed corresponding thereto in the first memory according to the target storage information associated with the target feature information, and send the data to be processed to the second processing device.
[0010] In some embodiments, the second memory is divided into multiple storage blocks, the first feature information of the stored data in the first memory and the associated storage information are stored in the multiple storage blocks according to a preset rule, the second memory is configured to use the second computing unit to compare second feature information of the retrieved data with third feature information corresponding to the storage blocks in the second memory, determine a target storage block according to the comparison result, and is configured to compare the second feature information of the retrieved data with the first feature information in the target storage block to obtain target feature information, where the third feature information corresponding to the storage block is determined according to the first feature information stored in the storage block.
[0011] In some embodiments, the multiple storage blocks are obtained by partitioning the second memory based on a target sorting result, where the target sorting result is obtained by sorting first feature information stored in the second memory based on a preset rule.
[0012] In some embodiments, the multiple storage blocks are obtained by acquiring the similarity between every two pieces of first feature information in the second memory and storing at least two pieces of first feature information with a similarity higher than a set threshold in the same storage block.
[0013] In some embodiments, the stored information includes a starting address and a storage length, where the storage length is determined according to the size of the stored data.
[0014] In some embodiments, the data to be processed received by the second processing device includes at least one of video data, audio data, distance data, and center-of-gravity data.
[0015] In some embodiments, the first processing device includes a central processing unit, the second processing device includes an artificial intelligence chip, and the computing unit includes an artificial intelligence computing unit.
[0016] In some embodiments, the first memory is a non-volatile memory, and the second memory is a random access memory.
[0017] In an embodiment of the present disclosure, a computing system includes a first processing device, a second processing device, and a storage device. The storage device includes a computing unit. The first processing device is configured to distribute computing tasks to the storage device and the second processing device. The storage device is configured to use the computing unit to perform computing processing on stored data in the storage device according to the received computing tasks to obtain intermediate processing results. The second processing device is configured to receive the data to be processed corresponding to the intermediate processing results sent by the storage device and perform computing processing on the data to be processed according to the received computing tasks to obtain target processing results, where the data to be processed includes a part of the stored data. By having the first processing device distribute tasks and having the computing unit in the storage device and the second processing device jointly complete the computing, the second processing device can perform computing processing on a part of the stored data to obtain the target processing results, thereby reducing data transfer, reducing power consumption, and improving processing speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary but non-limiting manner, where:
[0019] Figure 1A Schematically shows a schematic structural diagram of a computing system proposed according to an embodiment of the present disclosure;
[0020] Figure 1B Schematically shows a schematic structural diagram of another computing system proposed according to an embodiment of the present disclosure;
[0021] Figure 2 Schematically shows a schematic structural diagram of a second memory according to an embodiment of the present disclosure.
[0022] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed Embodiments
[0023] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and then implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0024] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, a device, an apparatus, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0025] In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0026] The principles and spirit of the present disclosure will be elaborated below with reference to several representative embodiments of the present disclosure.
[0027] The computing system proposed in the embodiments of the present disclosure is as Figure 1A shown. The computing system may include a first processing device 110, a second processing device 120, and a storage device 130. The storage device 130 includes a computing unit 140. Among them, the first processing device 110 may be a central processing unit CPU, and the second processing device 120 may be a processing chip. Taking the computing system as an AI computing system as an example, the second device 120 may be an AI chip. The storage device 130 may include non-volatile memories such as SSD or Flash Memory, and may also include random access memories such as DDR. Among them, the first processing device 110, the second processing device 120, and the storage device 130 are connected through a bus 100.
[0028] In this computing system, the first processing device 110 is used to distribute computing tasks to the second processing device 120 and the storage device 130.
[0029] By using the first processing device 110 to distribute computing tasks to the second processing device 120 and the storage device 130, the computing unit 140 in the storage device 130 can jointly complete the computing tasks with the second processing device 120. Those skilled in the art should understand that the number of the first processing device, the second processing device, and the storage device can be one or more, and the specific number of devices is not limited in this embodiment.
[0030] Generally, the computing unit 140 included in the storage device 130 has relatively weak computing power compared with the first processing device or the second processing device, and can be used to execute some specific computing tasks assigned by the first processing device. Taking the computing task as an AI computing task as an example, the computing task assigned to the computing unit in the storage device can be, for example, a cosine comparison task.
[0031] The storage device 130 is used to use the computing unit 140 to perform computing processing on the stored data in the storage device 130 according to the received computing task, and obtain an intermediate processing result.
[0032] Among them, the data stored in the storage device 130 is the original data processed by the computing system, which can be image data, audio files, etc. The computing unit 140 in the storage device 130 can perform computing processing on the locally stored data to obtain an intermediate processing result. Taking the computing task as an AI computing task as an example, the computing unit performs cosine comparison processing on the locally stored data to obtain a retrieval result as the intermediate processing result.
[0033] The second processing device 120 is used to receive the data to be processed corresponding to the intermediate processing result sent by the storage device 130, and perform computing processing on the data to be processed according to the received computing task to obtain a target processing result.
[0034] Among them, the storage device 130 can determine the corresponding data to be processed according to the intermediate processing result, where the data to be processed is a part of the stored data. That is to say, the storage device 130 does not send all the stored data to the second processing device 120, but sends the part of the stored data corresponding to the intermediate processing result to the second processing device 120. By discarding the stored data irrelevant to the computing task of the second processing device 120, data transmission can be reduced and system power consumption can be lowered.
[0035] In an embodiment of the present disclosure, a computing system includes a first processing device, a second processing device, and a storage device. The storage device includes a computing unit. The first processing device is configured to distribute computing tasks to the storage device and the second processing device. The storage device is configured to use the computing unit to perform computing processing on the stored data in the storage device according to the received computing tasks to obtain intermediate processing results. The second processing device is configured to receive the to-be-processed data corresponding to the intermediate processing results sent by the storage device and perform computing processing on the to-be-processed data according to the received computing tasks to obtain target processing results, where the to-be-processed data includes a part of the stored data. By having the first processing device distribute tasks and having the computing unit in the storage device and the second processing device jointly complete the computing, the second device can perform computing processing on a part of the stored data to obtain the target processing results, thereby reducing data transfer, lowering power consumption, and improving processing speed.
[0036] Figure 1B FIG. shows a computing system proposed in an embodiment of the present disclosure. As Figure 1B shown, the storage device 130 in the computing system includes a first memory 1301 and a second memory 1302. The first memory 1301 includes a first computing unit 1401, and the second memory 1302 includes a second computing unit 1402. Among them, the first memory 1301 can be a non-volatile memory, such as an SSD or Flash Memory. The second memory 1302 is a random access memory, such as DDR. In this case, the following method can be used to perform computing processing on the stored data in the storage device to obtain intermediate processing results.
[0037] Among them, the first memory 1301 is configured to use the first computing unit 1401 to perform computing processing on the stored data in the first memory 1301 to obtain a first processing result. The computing processing performed on the stored data is executed according to the computing tasks distributed by the first processing device 110, and the computing tasks are determined according to the computing power of the first computing unit 1401 in the first memory 1301 and the task cooperation among the various processing devices in the overall system.
[0038] The second memory 1302 is configured to receive and store the first to-be-processed data corresponding to the first processing result sent by the first memory 1301, where the first to-be-processed data includes a part of the stored data. The second memory 1302 is further configured to use the second computing unit 1402 to perform computing processing on the first to-be-processed data in the second memory 1302 to obtain the intermediate processing results.
[0039] In the embodiments of the present disclosure, data transmission can be performed between the first memory 1301 and the second memory 1302. For example, the first data to be processed corresponding to the first processing result can be sent to the second memory 1302. In related computing systems, it is usually necessary to move all stored data to the memory DDR for storage for the processing chip to perform computing processing. By moving a part of the stored data to the memory DDR in the embodiments of the present disclosure, data movement can be reduced and system power consumption can be lowered.
[0040] After receiving the first data to be processed from the first processing device 110, the second memory 1302 first stores the first data to be processed. After that, the second computing unit 1402 performs computing processing on the first data to be processed in the second memory 1302 to obtain the intermediate processing result. Since the second memory 1302 stores the first data to be processed, the second computing unit 1402 can directly perform computing on the first data to be processed stored locally to obtain the intermediate processing result.
[0041] In the embodiments of the present disclosure, taking the computing task as an AI computing task, all the original data is stored in the first memory 1301. The first computing unit 1401 performs preliminary computing and / or screening on the original data based on the allocated AI task, and sends the computing result and / or the corresponding screened data to the second memory 1302. The second computing unit 1402 performs further computing and / or further screening, and finally sends the computed and / or screened data to the AI chip. By performing computing and / or screening on the data level by level as described above, moving all the data is avoided.
[0042] In one example, a plurality of original images are stored in the first memory 1301. The first computing unit 1401 in the first memory 1301 first performs image recognition on the plurality of locally stored original images, and uses the recognition results that meet the set requirements as the first processing results. For example, "person, dog, house" recognized from the original images is used as the first processing result. Based on the recognition results of "person, dog, house", the first memory 1301 can obtain the images in the plurality of original images whose recognition results include "person, dog, house" as the first data to be processed, and send the first data to be processed to the second memory 120 for further AI processing. After receiving the first data to be processed, the second memory 120 can further perform recognition for each type of recognition result in the first processing result to obtain intermediate processing results. For example, after the first computing unit 1401 in the first memory 1301 recognizes "person" in the original image, the second computing unit 1402 in the second memory 1302 further performs recognition on the original image to obtain information such as "age, gender, height". For another example, body parts such as "hand, foot, face, ear" can be further recognized in the human body included in the original image. The intermediate processing results can be structured data, and the data to be processed and the intermediate processing results can be jointly fed back to the second processing device 120.
[0043] In some embodiments, the first computing unit 1401 in the first memory 1301 and the second computing unit 1402 in the second memory 1302 can be voice computing units, image computing units, or other types of computing units. The specific types of the computing units are not limited in the embodiments of the present disclosure.
[0044] Optionally, the first computing unit and the second computing unit can implement digital computing through a multiplier and implement analog computing through in-memory computing.
[0045] In some embodiments, the data to be processed received by the second processing device includes at least one of video data, audio data, distance data, and center of gravity data. Taking the second processing device as an AI chip as an example, the data to be processed received by the AI chip can be multi-dimensional data, and the AI chip makes intelligent decisions and judgments by integrating the multi-dimensional data.
[0046] In the case where the computing task of the computing system is a retrieval task, the computing system needs to retrieve all the stored data in the storage device to obtain the storage location of the target stored data that matches the retrieval data in the storage device, and obtain the target stored data. In some embodiments, the second memory 1302 may store index information of the stored data in the first memory 1301, where the index information includes information for indicating the storage location of the stored data in the first memory. The second memory 1302 uses the second computing unit 1402 to determine target index information that matches the retrieval data from the stored index information. For example, the second computing unit may be used to compare the information of the retrieval data with the index information, and obtain the target index information according to the comparison result, that is, using one or more index information that is closest to the information of the retrieval data as the target index information. Since the index information in the second memory is associated with the storage location of the stored data in the first memory, the storage location in the first memory can be determined through the target index information, and thus the data to be processed can be obtained.
[0047] In the embodiments of the present disclosure, by storing the index information of the stored data in the first memory in the second memory, the index information can be used to determine the data to be processed that matches the retrieval data in the first memory, reducing the amount of computation and improving the retrieval speed.
[0048] In some embodiments, the index information of the stored data in the second memory includes the first feature information of the stored data and the storage information of the stored data in the first memory, and the first feature information of the stored data is associated with the storage information.
[0049] Wherein, the stored data may be stored in units of storage objects. The storage objects may include various types such as people, animals, objects, virtual objects, etc., and the present disclosure does not limit the specific types of storage objects. The first feature information is obtained by performing feature extraction on at least one copy of the stored data corresponding to the storage object.
[0050] For example, the stored data of the storage object may be an image of a person, and each person corresponds to one or more images. By performing feature extraction on one or more images corresponding to each person, the first feature information corresponding to the storage object can be obtained.
[0051] For another example, the stored data of the storage object may be an audio file of a person, and each person corresponds to one or more audio files. By performing feature extraction on one or more audio files corresponding to each person, the first feature information corresponding to the storage object can be obtained.
[0052] That is, in the embodiments of the present disclosure, the first feature information corresponding to different stored data of the same storage object may be the same. For example, multiple different images of the same person may correspond to the same first feature information; for another example, images of the same person wearing different clothes or images of the same person from different angles may correspond to the same first feature information; for yet another example, images of multiple people with similar looks, such as images of twins or multiples, may correspond to the same first feature value.
[0053] The storage information is used to indicate the storage address of the stored data of the storage object, and may include the starting address and storage length of the stored data, where the storage length is determined according to the size of at least one piece of stored data corresponding to the storage object. In this case, the stored data in the storage device can be computationally processed by the following method to obtain an intermediate processing result.
[0054] First, use the second calculation unit in the second memory to compare the second feature information of the retrieved data with the first feature information of the stored data in the second memory, and determine the target feature information according to the comparison result. Wherein, the target feature information may be one or more pieces of first feature information that are closest to the second feature information of the retrieved data. Wherein, the second feature information is obtained by performing feature extraction on the retrieved data.
[0055] After that, use the first calculation unit in the first memory to obtain the data to be processed corresponding to the target feature information in the first memory according to the target storage information associated with the target feature information, and send the data to be processed to the second processing device. Since the first feature information of the stored data in the second memory is associated with the storage information of the stored data in the first memory, according to the target feature information, the storage information of the target data matching the retrieved data in the first memory can be determined, and the target feature information and the associated storage information can be used as intermediate results.
[0056] In the case where the first feature information of the storage object is associated with the storage information of at least one piece of stored data of the storage object, each storage object stored in the memory corresponds to stored first feature information and storage information, and the storage information indicates the storage address of the stored data of the storage object.
[0057] For example, in the case where multiple images of a person are stored in the memory, the first feature information of multiple images of this person is stored in the memory, and the first feature information is associated with the storage information of these multiple images.
[0058] By storing the first feature information of the storage object and the stored data in different regions while being associated with each other, it is possible to access the stored data using the first feature information as an index. In a data retrieval scenario, when the storage device is an SSD or flash memory, by using the first feature information as an index, some of the stored data can be filtered out by the local AI computing unit and transferred to the memory for comparison processing, and then the processed part of the stored data is sent to a CPU or an AI chip outside the storage device for further comparison processing to obtain the final retrieval result; compared with transferring all the stored data to the memory for comparison processing, it reduces data transfer and lowers the device power consumption. At the same time, storing the first feature information of the storage object and the stored data in different regions can also achieve the separation of the first feature information and the stored data, so that changes in the first feature information do not affect the stored data.
[0059] For example, in the case where the method for obtaining the first feature information of the storage object changes, such as the feature extraction algorithm is improved, resulting in an increase in the size of the feature map included in the first feature information. However, since the first feature information and the stored data are stored separately, the actual stored data will not change.
[0060] In the embodiments of the present disclosure, by storing the first feature information of the stored data and the storage information of the stored data in the first memory as index information in the second memory, and associating the first feature information of the stored data with the storage information of the stored data, it is possible to access the corresponding stored data using the first feature information as an index, avoiding transferring all the stored data to the memory for comparison processing, reducing data transfer, and improving the speed of data retrieval.
[0061] In the embodiments of the present disclosure, by storing the first feature information of the stored data in the first memory and the storage information of the stored data in the second memory, it is possible to access the corresponding stored data using the first feature information as an index, avoiding transferring all the stored data to the memory for comparison processing, reducing data transfer, and improving the speed of data retrieval.
[0062] In some embodiments, the second memory is divided into a plurality of storage blocks, and the first feature information of the stored data and the associated storage information in the first memory are stored in the plurality of storage blocks according to a preset rule. Wherein, each of the storage blocks corresponds to at least one storage object. The second memory is configured to use the second computing unit to compare the second feature information of the retrieved data with the third feature information corresponding to the storage blocks in the second memory, determine a target storage block according to the comparison result, and to compare the second feature information of the retrieved data with the first feature information in the target storage block to obtain target feature information, wherein the third feature information corresponding to the storage block is determined according to the first feature information stored in the storage block.
[0063] Among them, the set storage area in the second memory can be partitioned in a multi-level manner. For example, in the case where the first feature information of a plurality of storage objects is stored in the set storage area in the second memory, one or more of the plurality of storage objects can be divided into the same storage block to obtain a plurality of storage blocks at the first level. In the case where any storage block at the first level still contains a plurality of storage objects, further partitioning can still be performed in a similar manner to obtain a plurality of storage blocks at the second level, and so on.
[0064] In one example, the plurality of storage blocks are obtained by partitioning the second memory based on a target sorting result, where the target sorting result is obtained by sorting the first feature information stored in the second memory according to a preset rule. Specifically, the set storage area in the second memory can be partitioned in the following manner.
[0065] First, based on a preset rule, sort the first feature information of the plurality of storage objects stored in the set storage area.
[0066] The preset rule can be set according to the characteristics of the storage object itself. For example, when images of multiple people are stored in the memory, sorting can be performed according to the age of the people. For example, the first feature information corresponding to a person with a smaller age is stored more forward.
[0067] Sorting can also be performed according to other rules. For example, sorting can be performed according to the number of times the first feature information is searched. For example, the more times the first feature information of a storage object is searched, the more forward the first feature information of the storage object is arranged. Then, when applied to a data retrieval scenario, the first feature information will be compared earlier.
[0068] Next, divide the set storage in the second memory into a plurality of storage blocks according to the sorting result.
[0069] For the first feature information with similar sorting results, they can be stored in the same storage block. For example, in the case where the set storage area includes the first feature information of multiple storage objects, the first feature information of every n storage objects can be stored in the same storage block according to the sorting result of the first feature information.
[0070] In the embodiments of the present disclosure, by partitioning the set storage area according to the sorting result of the first feature information, storage objects with similar rankings of the first feature information can be stored in the same data block.
[0071] In some embodiments, the multiple storage blocks are obtained by acquiring the similarity between every two pieces of the first feature information in the second memory and storing at least two pieces of the first feature information with a similarity higher than a set threshold in the same storage block. Specifically, the set storage area can be partitioned according to the following method.
[0072] First, acquire the similarity between the first feature information of every two storage objects among the multiple storage objects stored in the set storage area.
[0073] Among them, the similarity between the first feature information of two storage objects can be determined according to the Euclidean distance between the feature vectors corresponding to the two pieces of the first feature information, or other methods can be used for calculation. The present disclosure does not limit the calculation method of the similarity.
[0074] Next, store the first feature information of at least two storage objects with a similarity higher than the set threshold in the same storage block.
[0075] Among them, for the storage objects stored in the same storage block, the similarity between the first feature information of every two storage objects can be higher than the set threshold, or the similarity between the first feature information of one storage object and at least one other storage object can be higher than the set threshold.
[0076] In the embodiments of the present disclosure, by partitioning the area according to the similarity between the first feature information of storage objects, storage objects with similar first feature information can be stored in the same data block.
[0077] Figure 2 Schematically shows a schematic structural diagram of a second memory according to an embodiment mode of the present disclosure, as Figure 2As shown, the first feature information of multiple storage objects is stored in a set storage area in the second memory. The stored data of the storage objects is stored in the first memory, and the first feature information of the storage objects is associated with the storage information of the stored data of the storage objects in the first memory. That is, the first feature information of the storage objects is stored in the second memory as index information. Among them, the first feature information of storage object 1 is feature value 1, the stored data of storage object 1 is picture 1, and feature value 1 is associated with the storage address 1 of picture 1; the first feature information of storage object 2 is feature value 2, the stored data of storage object 2 is picture 2a and picture 2b, and feature value 2 is associated with the storage address 2 of picture 2a and picture 2b; the first feature information of storage object 3 is feature value 3, the stored data of storage object 3 is picture 3, and feature value 3 is associated with the storage address 3 of picture 3. The set storage area 20 is divided into multiple storage blocks, which can also be called eigenvalue blocks, as Figure 2 shown, the feature value 1 of storage object 1 is stored in eigenvalue block A, and the feature value 2 of storage object 2 and the feature value 3 of storage object 3 are jointly stored in eigenvalue block B.
[0078] In some embodiments, the first feature information of the storage objects can be obtained by the following method.
[0079] First, use a pre-trained first feature extraction network to extract feature information from one of the stored data corresponding to the storage object to obtain sub-feature information.
[0080] Taking the stored data as an image as an example, a convolutional neural network can be used to extract the sub-feature information of the image.
[0081] After that, according to the sub-feature information of at least one of the stored data corresponding to the storage object, the first feature information of the storage object is obtained.
[0082] For example, the first feature information of the storage object can be obtained by concatenating the sub-feature information of each stored data; for another example, the first feature information of the storage object can be obtained by averaging or averaging and weighting the sub-feature information of each stored data.
[0083] When the first feature information of the stored data in the first memory and the associated storage information are stored in the multiple storage blocks according to a preset rule, the following method can be used to perform calculation processing on the stored data in the storage device to obtain an intermediate processing result.
[0084] First, use the second computing unit in the second memory to compare the second feature information of the retrieved data with the third feature information corresponding to the storage blocks in the second memory, and determine the target storage block according to the comparison result, where the third feature information corresponding to the storage block is determined according to the first feature information stored in the storage block. For example, one or more storage blocks closest to the second feature information can be determined as the target storage block.
[0085] After that, compare the second feature information of the retrieved data with the first feature information in the target storage block to obtain the target feature information.
[0086] Finally, use the first computing unit in the first memory to obtain the data to be processed corresponding to the target storage information associated with the target feature information in the first memory, and send the data to be processed to the second processing device.
[0087] In the embodiments of the present disclosure, by further comparing the second feature information in one or more storage blocks closest to the second feature information of the retrieved data, the comparison range can be reduced, the amount of data to be processed can be reduced, and the retrieval speed can be improved.
[0088] The following takes the image search by image scenario as an example to illustrate the data retrieval method proposed in the embodiments of the present disclosure. The system architecture applied by the method is as Figure 1B shown. The computing system includes a first processing device (CPU), a second processing device (AI chip), a first memory (SSD or Flash Memory), and a second memory (DDR), where the first memory includes a first computing unit (AI computing unit), and the second memory includes a second computing unit (AI computing unit). All images to be retrieved are stored in the first memory, and the data storage method in the second memory can refer to the description for Figure 2 which will not be elaborated here.
[0089] First, the CPU distributes AI tasks to the first memory, the second memory, and the AI chip, that is, instructs the first processing unit in the first memory, the second processing unit in the second memory, and the computing tasks to be executed by the AI chip.
[0090] The second computing unit compares the second feature information of the search image with the first feature information of the stored data in the second memory, where the first feature information and the second feature information are, for example, 1*512 feature data or 2*512 feature data. The second computing unit determines one or more first feature information closest to the second feature information according to the comparison result.
[0091] The first computing unit determines one or more corresponding original images in the first memory as data to be processed according to the determined one or more closest first feature information. For example, 10 images closest to the target image are obtained as data to be processed, and the data to be processed is sent to the AI chip. The AI chip directly compares and calculates the target image and the received data to be processed, and finally determines the image that matches the target image.
[0092] It should be noted that although several units / modules or sub-units / modules of the data storage device and the data retrieval device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple unit / modules.
[0093] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a device, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0094] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the data information processing device, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0095] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0096] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more of them in combination. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier to be executed by, or to control the operation of, data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information and transmit it to the appropriate receiver apparatus for execution by the data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0097] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform the functions by operating on input data and generating output. The processes and logical flows can also be performed by, or the apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0098] Suitable computers for executing computer programs include, by way of example, general and / or special purpose microprocessors, or any other type of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory and / or a random access memory. Basic components of a computer include a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, etc., or the computer will be operatively coupled to such mass storage devices to receive data from them or to transmit data to them, or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, such as a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning device (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.
[0099] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0100] Although this specification contains many specific implementation details, these should not be construed as limiting the scope of any invention or the scope of what is claimed, but rather as mainly describing the features of specific embodiments of a particular invention. Certain features that are described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may operate in certain combinations as described above and even be claimed as such initially, one or more features from a claimed combination may in some cases be removed from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0101] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various device modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and devices can generally be integrated together in a single software product or packaged into multiple software products.
[0102] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims may be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the drawings are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0103] The above description is only the preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope protected by one or more embodiments of this specification.
Claims
1. A computing system, characterized in that, It includes a first processing device, a second processing device, and a storage device. The storage device includes a computing unit. Among them, the first processing device is configured to distribute computing tasks to the storage device and the second processing device; the storage device is configured to use the computing unit to perform computing processing on the stored data in the storage device according to the received computing tasks, and obtain intermediate processing results; the second processing device is configured to receive the data to be processed corresponding to the intermediate processing results sent by the storage device, and perform computing processing on the data to be processed according to the received computing tasks to obtain target processing results. Among them, the data to be processed includes a part of the stored data. The first processing device, the second processing device, and the storage device are all connected to a bus, so that the first processing device, the second processing device, and the storage device can communicate through the bus; the storage device includes a first memory and a second memory. Among them, the first memory includes a first computing unit; the second memory includes a second computing unit; the first memory is configured to use the first computing unit to perform computing processing on the stored data in the first memory, and obtain first processing results; the second memory is configured to receive and store the first data to be processed corresponding to the first processing results sent by the first memory. Among them, the first data to be processed includes a part of the stored data; and is configured to use the second computing unit to perform computing processing on the first data to be processed in the second memory, and obtain the intermediate processing results; the storage device stores at least one storage object. The first feature information of each storage object and its corresponding stored data are stored in different areas of the storage device, and the first feature information corresponding to different stored data of the same storage object is the same.
2. The system according to claim 1, wherein index information of the stored data in the first memory is stored in the second memory; in the case where the computing task includes a retrieval task, the second memory is configured to use the second computing unit to determine target index information matching the retrieval data in the index information; the first memory is configured to use the first computing unit to obtain the data to be processed corresponding to the target index information in the first memory according to the target index information, and send the data to be processed to the second processing device.
3. The system according to claim 2, wherein the index information of the stored data in the second memory includes the first feature information of the stored data and the storage information of the stored data in the first memory, and the first feature information of the stored data is associated with the storage information; the second memory is configured to use the second computing unit to compare the second feature information of the retrieval data with the first feature information of the stored data in the second memory, and determine target feature information according to the comparison result; The first memory is configured to use the first computing unit to obtain corresponding data to be processed in the first memory according to the target storage information associated with the target feature information, and send the data to be processed to the second processing device.
4. The system according to claim 3, characterized in that, The second memory is divided into a plurality of storage blocks, and the first feature information of the stored data and the associated storage information in the first memory are stored in the plurality of storage blocks according to a preset rule. The second memory is configured to use the second computing unit to compare the second feature information of the retrieved data with the third feature information corresponding to the storage block in the second memory, determine the target storage block according to the comparison result, and compare the second feature information of the retrieved data with the first feature information in the target storage block to obtain the target feature information, where the third feature information corresponding to the storage block is determined according to the first feature information stored in the storage block.
5. The system according to claim 4, wherein The plurality of storage blocks are obtained by partitioning the second memory based on a target sorting result, where the target sorting result is obtained by sorting the first feature information stored in the second memory according to a preset rule.
6. The system according to claim 4, wherein The plurality of storage blocks are obtained by acquiring the similarity between every two pieces of first feature information in the second memory and storing at least two pieces of first feature information with a similarity higher than a set threshold in the same storage block.
7. The system according to any one of claims 3 to 6, characterized in that The storage information includes a start address and a storage length, where the storage length is determined according to the size of the stored data.
8. The system according to any one of claims 1 to 6, characterized in that, The data to be processed received by the second processing device includes at least one of video data, audio data, distance data, and centroid data.
9. The system according to any one of claims 1 to 6, characterized in that, The first processing device includes a central processing unit, the second processing device includes an artificial intelligence chip, and the computing unit includes an artificial intelligence computing unit.
10. The system according to any one of claims 1 to 6, characterized in that, The first memory is a non-volatile memory, and the second memory is a random access memory.
Citation Information
Patent Citations
Data storage method and device, computer equipment and readable storage medium
CN111224793A
Architecture and method for accelerating neural network calculation based on distributed weight storage
CN111275179A
Data storage method and data query method
CN111611418A