Embedded table storage method, storage device, electronic equipment, non-transitory computer readable storage medium and system
By storing and computing embedded tables in CXL-PNM devices, using their parallel lookup and near-storage computing characteristics, the memory constraints and storage constraints caused by excessive embedded tables are solved, and the performance and efficiency of the recommendation system are improved.
Patent Information
- Application Number
- CN202510031168.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively solve the memory constraints and storage constraints caused by excessive embedded tables. Especially in recommended systems, multi-GPU training will introduce network communication overhead, and the read delay and bandwidth of solid-state drives are poor, resulting in inefficient training.
By storing and computing embedded tables using CXL-PNM devices, the parallel search capability and near-storage computing characteristics of CXL-PNM devices can be used to reduce data transmission between the CPU and memory, free up the storage resources of the CPU, and reduce the operating burden of the CPU.
It improves the query efficiency of embedded tables, optimizes task execution efficiency, reduces the overall cost of ownership, solves the memory constraints and storage constraints caused by excessive embedded tables, and improves the performance of the recommended system.
Smart Images

Figure CN119987660A_ABST
Abstract
Description
Technical Field
[0001] Some example embodiments of the inventive concept relate to the field of storage and near memory computing, and in particular to a storage method, apparatus, electronic device, non-transitory storage medium, system and / or computer program product of an embedded table. Background Art
[0002] Embedding is a technology that decomposes concepts into feature vectors and / or automatically decomposes concepts (for example, software concepts, etc.) into feature vectors. It is used to map discrete features of input data to feature vectors for downstream neural network processing. Embedding plays a key role in recommendation systems. It can transform recommendation algorithms from "exact matching" to "fuzzy search", so that it can "draw inferences from one example" and improve the scalability of recommendation algorithms. Embedding Table (EMB) usually includes most of the parameters in the recommendation system model, and the size can reach TB level. The size of the embedding table is still rising rapidly, and it is difficult to put them into the memory of a single GPU (Graphics Processing Unit) for use.
[0003] Ideally, to solve larger-scale models, you only need to increase the number of GPUs, but the cost of GPUs is very high compared to CPUs, and multi-GPU training introduces additional network communication overhead, resulting in reduced training performance. In addition, the embedding table can be stored in a solid-state drive, but this solution has worse read latency and / or bandwidth, making training large-scale recommendation system models very time-consuming. Summary of the invention
[0004] The storage method and / or device of the embedded table, electronic device, non-transitory storage medium, system and / or computer program product provided in some example embodiments of the inventive concept can at least solve the above-mentioned technical problems and other technical problems not mentioned above.
[0005] According to at least one example embodiment of the inventive concept, a method for storing an embedded table is provided, the storage method comprising: obtaining an embedded table; storing the embedded table in a memory of at least one CXL-PNM device, the storing step enabling the at least one CXL-PNM device to perform a parallel lookup of the embedded table in the memory using a processing circuit of the at least one CXL-PNM device, wherein the at least one CXL-PNM device comprises the memory and the processing circuit.
[0006] Additionally, the storage method may include: generating the embedding table based on a data set including at least one transaction, wherein each data item in the data set is included in the transaction of the data set; storing the embedding vector corresponding to the item set in the embedding table into the memory in the at least one CXL-PNM device, and each item of the data set has an access relevance greater than an expected threshold.
[0007] Additionally, the storage method may include: obtaining an item set having an access relevance greater than a desired threshold based on a frequent pattern growth algorithm; and storing an embedding vector corresponding to the obtained item set in a memory of the at least one CXL-PNM device.
[0008] Additionally, the at least one CXL-PNM device is a plurality of CXL-PNM devices; the storage method may include: evenly distributing a computational task of obtaining a set of items in the data set whose access correlation is greater than an expected threshold to a processing circuit of each of the plurality of CXL-PNM devices; and completing the computational task by parallel operation of the processing circuit of each of the plurality of CXL-PNM devices.
[0009] Additionally, the storage method may include: generating a feature embedding table for each dimension of the data set, each of the feature embedding tables corresponding to the features of the corresponding dimension; using the processing circuit of the at least one CXL-PNM device to perform parallel calculations on each of the feature embedding tables, the parallel calculation step includes: generating a sub-frequent pattern tree corresponding to each of the feature embedding tables, setting a corresponding support threshold for each of the sub-frequent pattern trees, merging the attributes and links of each of the sub-frequent pattern trees to obtain a global frequent pattern tree, and sequentially processing the obtained global frequent pattern tree from the bottom layer of the obtained global frequent pattern tree to multiple node items included in the global frequent pattern tree, the sequential processing step includes: for a current node item among the multiple node items, based on the sub-frequent pattern tree corresponding to the current node item in the global frequent pattern tree, generating a conditional frequent pattern library corresponding to the current node item, based on the conditional frequent pattern library corresponding to the current node item and the support threshold of the sub-frequent pattern tree corresponding to the current node item, generating a conditional frequent pattern tree corresponding to the current node item, and based on the conditional frequent pattern tree corresponding to the current node item, obtaining an item set including the current node item whose access relevance is greater than an expected threshold.
[0010] According to at least one example embodiment of the inventive concept, there is provided a storage device for an embedded table, the storage device comprising: a first processing circuit configured to: obtain an embedded table; store the embedded table in a memory of at least one CXL-PNM device, the storing step causing the at least one CXL-PNM device to perform a parallel lookup of the embedded table in the memory using a second processing circuit of the at least one CXL-PNM device, wherein the at least one CXL-PNM device comprises the memory and the second processing circuit.
[0011] Additionally, the first processing circuit is further configured to: generate the embedding table based on a data set including at least one transaction, wherein each data item in the data set is included in the transaction of the data set; store the embedding vector corresponding to the item set in the embedding table into the memory in the at least one CXL-PNM device, and each item of the data set has an access relevance greater than an expected threshold.
[0012] Additionally, the first processing circuit is further configured to: obtain an item set whose access relevance is greater than an expected threshold based on a frequent pattern growth algorithm; and store the embedding vector corresponding to the obtained item set in the memory of the at least one CXL-PNM device.
[0013] Additionally, the at least one CXL-PNM device is a plurality of CXL-PNM devices, each of the plurality of CXL-PNM devices comprises the second processing circuit; the first processing circuit is further configured to: evenly distribute the computational task of obtaining a set of items in the data set whose access correlation is greater than an expected threshold to the second processing circuit of each of the plurality of CXL-PNM devices; and complete the computational task by the parallel operation of the second processing circuit of each of the plurality of CXL-PNM devices.
[0014] Additionally, the first processing circuit is further configured to: generate a feature embedding table for each dimension of the data set, each of the feature embedding tables corresponds to a feature of the corresponding dimension; use the second processing circuit of the at least one CXL-PNM device to perform parallel calculations on each of the feature embedding tables, the parallel calculation steps comprising: generating a sub-frequent pattern tree corresponding to each of the feature embedding tables, setting a corresponding support threshold for each of the sub-frequent pattern trees, merging the attributes and links of each of the sub-frequent pattern trees to obtain a global frequent pattern tree, and performing parallel calculations on the obtained global frequent pattern tree from the obtained A plurality of node items included in the global frequent pattern tree starting from the bottom layer of the global frequent pattern tree are processed sequentially, and the steps of sequential processing include: for a current node item among the plurality of node items, based on a sub-frequent pattern tree corresponding to the current node item in the global frequent pattern tree, generating a conditional pattern library corresponding to the current node item, based on the conditional frequent pattern library corresponding to the current node item and a support threshold of the sub-frequent pattern tree corresponding to the current node item, generating a conditional frequent pattern tree corresponding to the current node item, and based on the conditional frequent pattern tree corresponding to the current node item, obtaining an item set including the current node item whose access relevance is greater than an expected threshold.
[0015] According to at least one example embodiment, an electronic device includes: at least one processor; at least one memory configured to store computer executable instructions, wherein the at least one processor is configured to execute the computer executable instructions to perform the storage method of the embedded table as described above.
[0016] According to at least one example embodiment of the inventive concept, a non-transitory computer-readable storage medium storing computer-readable instructions is provided, wherein when the computer-readable instructions are executed by at least one processor, the at least one processor is caused to perform the storage method of the embedded table as described above.
[0017] According to at least one example embodiment of the inventive concept, a system including at least one computing device and at least one storage device storing computer-readable instructions is provided, wherein the computer-readable instructions implement the above-described method for storing an embedded table when executed by the at least one computing device.
[0018] According to at least one example embodiment of the inventive concept, a non-transitory computer-readable storage medium storing computer-readable instructions is provided, wherein the computer-readable instructions are executed by at least one processor to implement the above-mentioned method for storing an embedded table. One or more technical solutions provided in the example embodiments of the inventive concept bring at least the following beneficial effects:
[0019] According to at least one example embodiment of the inventive concept, a storage method and / or apparatus for an embedded table, an electronic device, a non-transitory storage medium, a system and / or a computer program product may include a CXL-PNM device (a PNM device based on CXL extension, CXL: Compute Express Link, PNM: Processing near Memory), which may be used to store and / or calculate the embedded table, so as to solve the problems of memory limitation and storage constraint caused by the embedded table being too large through storage extension based on the CXL-PNM device, etc., and when used in a recommendation system, the performance of the recommendation system may be improved. Additionally, through the near storage calculation of the CXL-PNM device, the data transmission between the computing modules (for example, taking the CPU (Central Processing Unit, central processing unit) as an example) and / or the memory resource utilization may be reduced or greatly reduced, and / or the storage resources of the CPU may be released, and the operation and / or computing burden of the CPU may be reduced, so as to improve the execution efficiency of the recommendation system and / or improve and / or optimize the performance of the recommendation system, etc.
[0020] In addition, the embedding vectors corresponding to the itemsets with high access correlation are centrally stored in the CXL-PNM device, which can improve the query efficiency of the embedding vectors.
[0021] In addition, evenly distributing computing tasks among multiple CXL-PNM devices can improve and / or optimize task execution efficiency, etc.
[0022] In addition, through the multi-support-Parallel Frequent Pattern Growth algorithm (MS-PFP-Growth), it is possible to more efficiently mine item sets with high access correlation in data items based on the storage and / or computing architecture of multiple CXL-PNM devices, so that the embedding vectors corresponding to the item sets with high access correlation can be more efficiently stored in multiple CXL-PNM devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings herein are incorporated in and include a part of the specification, illustrate some example embodiments consistent with the inventive concept, and together with the description are used to explain the principles of one or more example embodiments of the inventive concept, and shall not unduly limit the inventive concept.
[0024] Figure 1 A flowchart illustrating a storage method of an embedded table in at least one example embodiment of the inventive concept.
[0025] Figure 2A A schematic diagram of a memory expansion structure based on CXL in the related art is shown, and Figure 2B A schematic diagram of the structure of memory expansion based on CXL-PNM according to at least one example embodiment of the inventive concept is shown.
[0026] Figure 3A A schematic diagram of a SSD-based computing storage architecture in the related art is shown, and Figure 3B A schematic diagram of a CXL-PNM-based computational storage architecture according to at least one example embodiment of the inventive concepts is shown.
[0027] Figure 4 A schematic diagram showing a comparison of the access hit rates of embedded vectors between a computational storage architecture in the related art and a computational storage architecture in at least one example embodiment of the inventive concept.
[0028] Figure 5 A framework diagram showing a computational storage architecture based on CXL-PNM in at least one example embodiment of the inventive concept.
[0029] Figure 6 A schematic diagram illustrating a multiple support concept based on an FP-growth algorithm in at least one example embodiment of the inventive concept.
[0030] Figure 7 A schematic flow chart showing an MS-PFP-Growth algorithm in at least one example embodiment of the inventive concept.
[0031] Figure 8 A schematic diagram illustrating a frequently accessed item set generated in at least one example embodiment of the inventive concept.
[0032] Fig. 9 A schematic diagram showing the structure of a header table in at least one example embodiment of the inventive concept.
[0033] Fig.10 A schematic diagram illustrating a global frequent pattern tree in at least one example embodiment of the inventive concept.
[0034] Fig.11 A schematic diagram of a process of generating a conditional frequent pattern library corresponding to node I4 in at least one example embodiment of the inventive concept is shown.
[0035] Fig.12 A schematic diagram illustrating storing highly accessed related item sets in the form of a conditional frequent pattern tree in at least one example embodiment of the inventive concept.
[0036] Fig.13 A block diagram illustrating a storage device embedded with a table in an exemplary embodiment of the inventive concept.
[0037] Fig.14 A block diagram of an electronic device illustrating at least one example embodiment of the inventive concept. DETAILED DESCRIPTION
[0038] In order to enable ordinary persons in the art to better understand the technical solutions of the exemplary embodiments of the inventive concept, the technical solutions in some exemplary embodiments of the inventive concept will be clearly and completely described below with reference to the accompanying drawings.
[0039] It should be noted that the terms "first", "second", etc. in the descriptions and claims of some example embodiments of the inventive concept and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the example embodiments of the inventive concept described here can be implemented in an order other than those illustrated and / or described here. The implementations described in the following example embodiments do not represent all implementations consistent with the example embodiments of the inventive concept. Instead, they are merely examples of some aspects of the inventive concept as detailed in the appended claims.
[0040] It should be noted that the phrase "at least one of the items" mentioned here means "any one of the items", "a combination of any number of the items" and / or "all of the items". For example, "including at least one of A and B" means the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" means the following three parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.
[0041] As the amount of information on the Internet continues to grow exponentially, users are facing a serious problem of information overload. Therefore, recommendation systems came into being. Recommendation systems aim to meet users' growing personalized needs through information filtering, thereby helping users to make information choices more quickly and / or efficiently. User behavior data in recommendation systems usually present a long-tail distribution, and most items (e.g., or data items, referring to recommended objects such as items in user behavior data, etc.) do not exist and / or there are very few user behavior data. Therefore, recommendation algorithms cannot be satisfied with just remembering "common high-frequency" patterns, but also need to be improved by being able to mine and / or automatically mine "low-frequency long-tail" patterns.
[0042] Traditional recommendation algorithms can memorize high-frequency patterns but cannot scale. Therefore, embedding tables can be used to improve the scalability and / or efficiency of recommendation algorithms, and to mine and / or automatically mine low-frequency, long-tail and / or niche patterns, thereby achieving personalized recommendations, etc.
[0043] Since the embedding table used in the recommendation system occupies a large amount of storage space, two solutions have been proposed to address the problem of limited storage of the embedding table.
[0044] One proposed solution is to use a multi-GPU storage system such as HugeCTR. HugeCTR is an open source recommendation system framework that can split the embedding table to multiple GPU cards on multiple nodes for calculations, where each GPU card processes a portion of the data. However, due to the amount of network communication overhead required to facilitate the splitting of the embedding table, this will cause huge and / or large performance waste. The efficiency experiment of HugeCTR's distributed training with different GPUs connected by WDL (Wireless Data Link) showed that up to 90% of the training time will be spent on obtaining and updating the embedding parameters, and also showed that data transfer dominates the training cycle. With the emergence of more powerful accelerators and the slow growth of network bandwidth, the gap between data processing speed and data transmission speed will become larger and larger, and communication bottlenecks will become a more serious problem.
[0045] Another solution is to offload the embedding table to other devices. For example, the embedding table can be stored in an SSD. However, since different data corresponds to different embedding vectors, and as the data set and embedding table increase, this method will greatly increase the access error rate of the embedding vector, and the possibility of saving the correct embedding vector in HBM (High Bandwidth Memory) / DRAM (Dynamic Random Access Memory) will also decrease. In addition, the capacity of SSDs can reach 10TB or even hundreds of TB, but it takes several Gb / s to read data from the SSD to the CPU memory, so the bandwidth requirement is high and the data access speed is slow. Therefore, it is not appropriate to store the entire embedding table on the SSD because the reading overhead will be very large.
[0046] RecShard is a fine-grained embedding table partitioning and placement technology for Deep Learning Recommendation Models (DLRMs). This technology is based on the characteristics of EMB and has the hierarchical nature of memory architecture, improving performance and efficiency by improving memory usage and reducing access latency.
[0047] RecShard can use a multi-GPU system to store the embedding table. The algorithm operates as follows: (1) randomly sample 1% of the entire data set; (2) find the embedding vector of the sampled result; (3) distribute the embedding vector to different GPU HBMs. This algorithm can reduce the GPU error rate by pre-fetching some embedding vectors into HMB, but there are three problems: (1) it occupies too many CPU resources; (2) it occupies too many GPU HBMs; (3) it provides a solution that does not update the embedding vector during training, so the error rate of the embedding vector is still high.
[0048] RecSSD is a solid-state drive storage system based on approximate data processing, customized for neural recommendation inference. RecSSD stores frequently accessed embedding tables in the host's DRAM and infrequently accessed embedding tables in RecSSD. RecSSD can use internal bandwidth to perform "lookups" in embedding tables, implement vector operations, and return results in the form of logical blocks. However, RecSSD will take up too much host memory space (for example, increased use of the host memory monitor, etc.), and performing vector operations in FTL (Flash Translation Layer) will increase and / or cause unnecessary GPU waits, delays, and / or delays. In addition, only vector addition and vector concatenation can be implemented in FTL, so the application scenarios of RecSSD are very limited, and implementing complex vector operations will result in longer GPU wait times.
[0049] In summary, the two solutions to the problem of limited embedded table storage in the related art have at least the following problems: (1) the total cost of ownership (TCO) of the multi-GPU solution is high; (2) increased and / or excessive network communication overhead causes increased and / or higher performance waste; and / or (3) high access error rate causes large additional read overhead.
[0050] In order to solve one or more of the above problems, one or more exemplary embodiments of the inventive concept provide a storage method and / or device for storing an embedded table, an electronic device, a non-transitory storage medium and / or a system for storing a computer program product. The storage method of the embedded table can be executed by a host (e.g., a host device, etc.), and the embedded table can be stored and / or calculated using a CXL-PNM device to solve the problems of memory limitation and / or storage constraints caused by the embedded table being too large through storage expansion based on the CXL-PNM device. When used in a recommendation system, the performance of the recommendation system can be improved; and through the near storage calculation of the CXL-PNM device, the data transmission between the CPU and the memory can be reduced or greatly reduced, the storage resources of the CPU can be released, and / or the operating burden of the CPU can be reduced, thereby improving the execution efficiency of the recommendation system and / or enhancing and / or optimizing the performance of the recommendation system.
[0051] Next, we will refer to Figures 1 to 14 Some example embodiments of the inventive concepts are described in detail.
[0052] Figure 1 A flowchart illustrating a storage method of an embedded table in at least one example embodiment of the inventive concept.
[0053] Reference Figure 1 , in operation 101, an embedding table may be obtained.
[0054] According to at least one example embodiment of the inventive concept, an embedding table used in fields such as recommendation systems, image classification, natural language processing, etc. can be generated and / or obtained. For example, the embedding table can be generated and / or learned during artificial intelligence, machine, learning, and / or neural network training, so different tasks and data sets may expect and / or need to generate and / or obtain different embedding tables. In at least one example embodiment, a pre-trained embedding table can be obtained to speed up the training process and / or improve performance, but the example embodiments are not limited to this.
[0055] At operation 102, the embedding table may be stored in a memory of at least one CXL-PNM device, so that during the embedding table lookup, a parallel lookup is performed in the memory by a computing module (e.g., a processing circuit, a computing circuit, a computing device, etc.) of at least one CXL-PNM device, wherein the at least one CXL-PNM device includes or comprises a memory and a computing module, but is not limited thereto. According to some example embodiments, the computing module may be implemented as a processing circuit. The processing circuit may include hardware or a hardware circuit including a logic circuit; a hardware / software combination, such as a processor that executes software and / or firmware; or a combination of the two. For example, more specifically, the processing circuit may include, but is not limited to, a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a system on a chip (SoC), a programmable logic unit, a microprocessor, an application specific integrated circuit (ASIC), etc., but is not limited thereto.
[0056] According to at least one example embodiment of the inventive concept, the recommendation system may include a new CXL-PNM-based computing storage architecture to solve the memory limitation problem caused by, for example, an overly large embedding table, but is not limited thereto. It is understood that the CXL-PNM-based computing storage architecture in at least one example embodiment of the inventive concept may not be limited to being applied to the training, verification, and use of the model, and the training process is taken as an example below.
[0057] Processing Near Memory (PNM) is a technology that integrates memory and logic chips (e.g., circuit systems, etc.) into advanced integrated circuit packages, but example embodiments are not limited to this. By utilizing memory for data calculations, data movement between the CPU and / or GPU and memory can be reduced. Computing functions are performed closer to the memory, which can reduce bottlenecks in data transmission between the CPU and / or GPU and memory, etc. The open standard CXL can be used in conjunction with PNM to facilitate the expansion of memory capacity. Relevant tests have shown that PNM solutions based on the CXL interface can improve or significantly improve the performance of applications such as recommendation systems and / or in-memory databases that expect and / or require high memory bandwidth.
[0058] CXL TM (Compute Express Link TM , Compute Express Link) is an open standard for high-speed processor-to-device and processor-to-memory interfaces that enable more efficient use of memory and accelerators used with processors. CXL technology can interconnect disparate compute and storage resources to improve system performance and / or efficiency. CXL TMCan be used in conjunction with other technologies, such as PNM, to help facilitate memory capacity expansion.
[0059] The CXL-PNM device is a device that integrates arithmetic functions into memory semiconductors. By placing the computing functions next to the memory, the data movement between the CPU and / or GPU and the memory is reduced, thereby reducing bottlenecks and maximizing the processing power of the CPU.
[0060] Figure 2A A schematic diagram of a memory expansion structure based on CXL in the related art is shown, and Figure 2B A schematic diagram of the structure of memory expansion based on CXL-PNM according to at least one example embodiment of the inventive concept is shown.
[0061] Reference Figure 2A , shows the memory expansion based on CXL in the related art. The CPU and the device memory of CXL are connected through the CXL controller (CXL controller, CXL CTRL) to achieve high-speed, low-latency, high-bandwidth data transmission and sharing between computing resources such as CPU and storage resources such as CXL device memory.
[0062] Reference Figure 2B , shows a memory expansion based on CXL-PNM of at least one example embodiment of the inventive concept, and the memory expansion based on CXL-PNM combines CXL and PNM technologies through a CXL controller to implement CXL-PNM memory expansion based on a CXL controller. CXL technology can provide one or more high-speed and / or low-latency data transmission channels, so that the CPU and the accelerator in the PNM can efficiently share and / or access memory resources, etc. In CXL technology, there are mainly the following key components and concepts:
[0063] CXL CTRL: This is the control device and / or circuitry of the CXL, which may be responsible for managing and / or controlling CXL connections and / or communications, but is not limited to these. It may handle protocols and / or signals associated with the CXL, ensuring that communication between devices is correct.
[0064] DOR (Device On-Ramp): DOR is a device access point that allows devices (such as accelerators, memory buffers, etc.) to connect to the CXL network. DOR can provide a standardized interface for devices to communicate with other components in the CXL network.
[0065] MC (Memory Controller): MC is the component responsible for managing memory access. In a CXL environment, MC can implement and / or ensure that access to shared memory is efficient and / or consistent, but is not limited to this. In CXL technology, MC can work in conjunction with the host CPU, or with accelerator devices (such as PNM, etc.), and / or move to memory buffer chips, etc.
[0066] CXL.Jo: CXL.Jo is a joint interface of CXL, which can allow two or more CXL devices to work with each other and, for example, share memory and / or other resources. Through CXL.Jo, the devices can operate as a whole (e.g., a single entity, etc.), thereby improving overall performance and / or efficiency, etc.
[0067] CXL.mem: CXL.mem is a memory protocol of the CXL protocol and may define a transmission interface between the CPU and the memory to implement and / or ensure that the CPU may efficiently access and use the shared memory in the CXL network, etc.
[0068] Device Memory: Device memory can refer to memory on the accelerator and / or other devices (e.g., host, etc.). In the CXL architecture, these memories can be shared and / or accessed by the CPU and / or other devices (e.g., PNM, etc.), thereby improving overall performance and / or resource utilization, etc.
[0069] PNM solutions based on the CXL interface can provide significantly higher or better performance in applications where high memory bandwidth is expected and / or required.
[0070] It is understandable that PNM can also be dispersed into memory to implement a PNM solution based on a CXL interface, etc.
[0071] According to at least one example embodiment of the inventive concept, the calculation and storage of the embedded table can be placed in the CXL-PNM device, but is not limited to this. For example, operations of embedded vectors (referred to as vector operations) include but are not limited to vector concatenation, vector summation, vector multiplication and / or vector average, etc., which can all be placed on the PNM and can be processed using PIM (Processing-in-Memory) logic, but the example embodiment is not limited to this. For example, the embedded table can be unloaded to multiple CXL-PNM devices, etc. In the case where the memory capacity of the CXL-PNM is, for example, 512GB, the number of CXL-PNM devices can be set from 1 to N according to the expectation and / or requirement of the embedded table storage space, N is an integer greater than 1, but the example embodiment is not limited to this. Assuming that the EMB capacity is M (GB), the number of CXL-PNM devices required is N=ceil(M / 512)(GB). It can be understood that ceil (e.g., rounding function) means rounding up to the next integer.
[0072] Figure 3A A schematic diagram of a SSD-based computing storage architecture in the related art is shown, and Figure 3B A schematic diagram of a computational storage architecture based on CXL-PNM in at least one example embodiment according to the inventive concepts is shown.
[0073] Reference Figure 3A In the related art, the embedded table (EMB Table) can be stored in the SSD (Solid State Disk or Solid State Drive). During the training process, the operations related to the embedded table can mainly include the following items: ① The embedded table can be unloaded (EMB offload) by the CPU to store the acquired embedded table in the SSD; ② The CPU performs an embedded table lookup (EMB lookup) based on and / or based on the training data to search in the embedded table stored in the SSD; ③ Through the embedded table cache mechanism (EMB cache) in the CPU, the embedded vector found in the embedded table based on and / or based on the training data is read from the SSD to the DRAM of the CPU; ④ The processing circuit (Compute) for vector operation (Vector operation) is implemented by the GPU to perform vector operations on the embedded vectors read into the DRAM of the CPU.
[0074] It can be seen that this method of the related art will perform a large amount of frequent data transmission between the CPU and the memory (SSD) storing the embedded table, and between the CPU and the GPU. In addition, this multi-GPU implementation is not only costly but also wastes computing resources.
[0075] Reference Figure 3B In at least one example embodiment of the inventive concept, the embedding table can be stored by the CXL-PNM device. During the training process, the operations related to the embedding table can mainly include the following items: ① The CPU can be used to perform unloading of the embedding table to store the acquired embedding table in the memory of the CXL-PNM device. The memory of the CXL-PNM device mentioned in one or more example embodiments of the inventive concept can be a memory (Memory (CXL)) expanded through the CXL interface; ② The CPU performs a search for the embedding table according to and / or based on the training data to search in the embedding table stored in the CXL-PNM device; and / or ③ The PNM in the CXL-PNM device implements a processing circuit for vector operations to perform vector operations on the embedded vector found from the memory of the CXL-PNM device, etc., but the example embodiments are not limited thereto.
[0076] It can be seen that this approach can reduce or significantly reduce the data transmission (e.g., the amount of data transmitted, etc.) between the CPU and the memory storing the embedding table (the memory of the CXL-PNM device), without increasing the data transmission burden between the CPU and the GPU, and / or without requiring multi-GPU implementation, thereby reducing costs and / or saving computing resources, etc. Integrating the computing module for vector operations with the memory storing the embedding table can also further reduce the data transmission consumption in the training process, etc.
[0077] It is understandable that the embedding table is a special key-value mapping, the "key" is usually the feature ID (and / or index) in the original data, and the "value" is the low-dimensional vector corresponding to the feature ID, but the exemplary embodiment is not limited thereto. This low-dimensional vector is usually obtained through model training and can capture the potential relationship between different features in the original data.
[0078] It is understandable that the CPU can be used to execute relevant computer-readable instructions and / or algorithms such as unloading, searching, etc.; this includes reading the location information of the embedded table, allocating system resources (such as memory, etc.) to temporarily store data and / or perform data migration, deletion, etc.; and can quickly locate the target data through traversal, indexing and / or other data structures, which usually involves reading and / or comparing operations on the data in the table, etc., but the example embodiments are not limited to this.
[0079] As the main memory of the embedded system, DRAM can be used to store the operating system, application programs and / or runtime data. During the process of unloading the embedded table, DRAM can be used to temporarily store data read from the table, store temporary data and / or logs generated during the unloading process, etc. The fast access characteristics of DRAM can enable these operations to be performed efficiently.
[0080] According to at least one example embodiment of the inventive concept, when there are multiple CXL-PNM devices, during the embedded table lookup, the processing circuits of the multiple CXL-PNM devices may perform parallel lookups in the memory to improve the lookup efficiency, but example embodiments are not limited thereto.
[0081] According to at least one example embodiment of the inventive concept, the operation of storing an embedding table in the memory of at least one CXL-PNM device may include, but is not limited to: storing the embedding vectors corresponding to the item sets in the embedding table whose access relevance is greater than an expected and / or preset threshold in the memory of the same CXL-PNM device in at least one CXL-PNM device, wherein the embedding table may be constructed based on a data set including at least one transaction, and each data item in the item set (e.g., a data set) may be included in the transaction of the data set, but is not limited thereto.
[0082] According to at least one example embodiment of the inventive concept, an embedded table may be generated for any one or more data sets such as user behavior data, social media data, transaction data, medical record data, etc., and each data set may contain at least one corresponding transaction data.
[0083] According to at least one example embodiment of the inventive concept, the item sets in the data set whose access relevance is greater than the expected and / or preset threshold (for example, high access relevance item sets, etc.) can be calculated through association rule mining algorithms such as Apriori (prior) and Eclat (Equivalence Class Transformation), and the embedding vectors corresponding to the high access relevance item sets in the embedding table are centrally stored in the memory of the same CXL-PNM device to achieve centralized access during the training process and improve the query efficiency of the embedding vectors.
[0084] Figure 4 (a) is a schematic diagram showing a comparison of access hit rates of embedded vectors of computational storage architectures in related arts, and (b) is a schematic diagram showing a computational storage architecture in at least one example embodiment of the inventive concept.
[0085] Reference Figure 4(a), in the related art, during model training, a portion of the embedding table is pre-saved in the HBM of the GPU through embedding table copy (EMB Copy). For the embedded vectors saved in the GPU, if they are accessed (and / or read) using embedding search during training, they can be directly hit by key matching (for example, hit keys), while the embedded vectors that are not saved in the GPU will not be hit in the GPU when accessed (for example, miss keys), and need to be read from the CPU. The embedded table stored in the DRAM of the CPU may also be obtained from a storage device such as an SSD through embedding table copy. Therefore, in the process of embedding table lookup, the related art at least needs to perform operations such as embedding access (and / or item access, Items Access) and embedding table copy through the index corresponding to the item (Item), and the hit rate of the first access is low.
[0086] Reference Figure 4 (b) According to at least one example embodiment of the inventive concept, if the embedded table is stored in the memory of CXL, it can be directly read through PNM without the situation of access failure. Moreover, during the embedded table lookup, only operations such as item access can be performed without the need to perform embedded table copy operations.
[0087] According to at least one example embodiment of the inventive concept, the step of storing the embedding vectors corresponding to the item sets in the embedding table whose access relevance is greater than the expected and / or preset threshold in the memory of the same CXL-PNM device in at least one CXL-PNM device may include but is not limited to: the item sets whose access relevance is greater than the expected and / or preset threshold may be obtained based on the frequent pattern growth algorithm through parallel calculation of the calculation modules of at least one CXL-PNM device; and / or the embedding vectors corresponding to the item sets whose access relevance is greater than the expected and / or preset threshold may be stored in the memory of the same CXL-PNM device, but is not limited to this.
[0088] According to at least one example embodiment of the inventive concept, FP-growth (Frequent Pattern Growth) is an efficient correlation rule mining algorithm that can be applied to concurrent computing of big data, but the example embodiments are not limited thereto.
[0089] Figure 5 A framework diagram showing a computational storage architecture based on CXL-PNM in at least one example embodiment of the inventive concept.
[0090] See also Figure 5According to at least one example embodiment of the inventive concept, the process of training a model to search, operate and / or update an embedding table may include a host side (e.g., at least one host device, etc.) and a CXL-PNM device side, but the example embodiment is not limited thereto. The CXL-PNM device side may include multiple CXL-PNM devices to increase storage capacity; the embedding table may be split and stored in multiple CXL-PNM devices. The splitting method may be to store the embedding vectors corresponding to each high-access related item set mined by the frequent pattern growth algorithm in the memory of at least one CXL-PNM device; therefore, in the memory of the CXL-PNM device, the embedding vector is stored in the form of a conditional frequent pattern tree (Conditional FP-Tree), that is, each item in the conditional frequent pattern tree is a high-access related item set, and the index corresponding to the high-access related item set is stored in the memory of one or more CXL-PNM devices, but the example embodiment is not limited thereto.
[0091] It can be understood that the item corresponds to the index of the embedding vector, but is not limited to this.
[0092] The host side (e.g., a host device, etc.) may generate a corresponding embedding table lookup instruction according to and / or based on the data set, so that the CXL-PNM device side may search for the corresponding embedding vector in the memory of the CXL-PNM device according to and / or based on the embedding table lookup instruction, and may also perform vector operations on the found embedding vector in the processing circuit (PNM) of the CXL-PNM device. It is understandable that vector operations are not necessary operations for each embedding table lookup, or in other words, vector operations are optional operations. Then, the host side reads the found embedding vector from the CXL-PNM device side, that is, the feature vector (Feature Vector) corresponding to the current data set, and may use the feature vector for the training process of the neural network, etc.
[0093] According to at least one example embodiment of the inventive concept, the operation of obtaining an item set whose access relevance is greater than an expected and / or preset threshold value may include but is not limited to: evenly distributing the computational task of obtaining an item set whose access relevance is greater than an expected and / or preset threshold value in a data set to the computational modules of each CXL-PNM device, so as to complete the computational task through the parallel work of the computational modules of each CXL-PNM device, but is not limited thereto. Evenly distributing the computational task of mining high access relevance item sets among each CXL-PNM device can further improve computational efficiency.
[0094] According to at least one example embodiment of the inventive concept, the step of obtaining an item set whose access relevance is greater than an expected and / or preset threshold value may include but is not limited to: a feature embedding table corresponding to the feature of each dimension of the data set may be constructed respectively; and / or each feature embedding table may be calculated in parallel by a computing module of at least one CXL-PNM device to perform the following operations: a sub-frequent pattern tree corresponding to each feature embedding table may be constructed, and a corresponding support threshold may be set for each sub-frequent pattern tree; the attributes and links of each sub-frequent pattern tree may be merged to obtain a global frequent pattern tree; the obtained global frequent pattern tree may be processed sequentially from the bottom node items, and for the current node item, a conditional frequent pattern library corresponding to the current node item may be constructed according to and / or based on the sub-frequent pattern tree corresponding to the current node item in the global frequent pattern tree; a conditional frequent pattern tree corresponding to the current node item may be constructed according to and / or based on the conditional frequent pattern library corresponding to the current node item and the support threshold of the sub-frequent pattern tree corresponding to the current node item; and / or an item set including the current node item whose access relevance is greater than an expected and / or preset threshold may be obtained according to the conditional frequent pattern tree corresponding to the current node item, etc., but the example embodiments are not limited thereto.
[0095] According to at least one example embodiment of the inventive concept, the minimum support (Min_sup) of FP-growth determines the overall performance of the algorithm. If the minimum support is not set properly, a large amount of "long tail" data will be filtered, which will have an adverse effect on the performance of the recommendation system. Therefore, this problem can be overcome by the concept of multiple support based on the FP-growth algorithm in at least one example embodiment of the inventive concept.
[0096] Figure 6 A schematic diagram illustrating a multiple support concept based on an FP-growth algorithm in at least one example embodiment of the inventive concept.
[0097] Reference Figure 6 , a feature embedding table (FeatureEMB Table) can be constructed and / or generated for the features of each dimension of the data set; in other words, the embedding table can be split according to and / or based on the features of each dimension of the corresponding data set to obtain feature embedding tables corresponding to the features of each dimension. Then, a corresponding sub-frequent pattern tree (Sub FP-Tree) can be constructed and / or generated based on each feature embedding table. For each sub-frequent pattern tree, different minimum supports can be maintained according to and / or based on the access information of different feature embedding tables, thereby improving and / or optimizing the algorithm performance, but the example embodiments are not limited to this.
[0098] According to at least one example embodiment of the inventive concept, when calculating the minimum support, different support thresholds can be set for all feature embedding tables according to and / or based on the number of items (EMB items) contained in the feature embedding table stored in the CXL device, the correlation between items, and the access popularity of the feature embedding table. A smaller support threshold can be set for feature embedding tables that are accessed more frequently. The larger the support threshold, the more accurate the relevant items obtained, but the fewer the number. Just as an example, if the support threshold of feature embedding table 1 (EMB table 1) is set to 50%, then the minimum support = 0.5 × the total number of access IDs, and the value can be rounded down to a decimal point.
[0099] According to at least one example embodiment of the inventive concepts, the multiple support concept based on the FP-growth algorithm may be implemented as an MS-PFP-Growth algorithm, but example embodiments are not limited thereto.
[0100] Figure 7 A schematic flow chart showing an MS-PFP-Growth algorithm in at least one example embodiment of the inventive concept.
[0101] Reference Figure 7 For example, the workflow of the MS-PFP-Growth algorithm can be as follows: (1) a sub-frequent pattern tree can be generated: the embedding table can be split into multiple feature embedding tables and stored in the memories of multiple CXL-PNM devices, and then the computing modules of each CXL-PNM device can calculate all feature embedding tables in parallel to generate a sub-frequent pattern tree corresponding to each feature embedding table, each sub-frequent pattern tree having a corresponding minimum support corresponding to the support threshold; (2) sub-frequent pattern trees can be merged: after all sub-frequent pattern trees are generated, they can be merged to generate a global frequent pattern tree; (3) a conditional frequent pattern tree can be generated: the global frequent pattern tree in (2) can be used to generate a conditional frequent pattern tree; and / or (4) a highly accessed related item set can be mined: the conditional frequent pattern tree can be recursively mined while growing the set of highly accessed related items it contains, so as to mine a highly accessed related item set according to and / or the conditional frequent pattern tree; if the conditional frequent pattern tree contains only one path, the contained highly accessed related item set can be directly generated, and the embedding vectors (e.g., EMB) corresponding to each item in the highly accessed related item set in the embedding table can be directly generated. Items) are stored in a CXL extended memory, but the exemplary embodiment is not limited thereto. In the process of mining highly visited related item sets, the support threshold corresponding to each sub-frequent pattern tree may be used.
[0102] According to at least one example embodiment of the inventive concept, a sub-frequent pattern tree can be generated in the following manner: it can be assumed that: I = {a_1, a_2, ..., a_m} is a set of m items (e.g., projects) contained in a feature embedding table; D = {I_1, I_2, ..., I_n} is data containing n access item sets, I_i = {a_i1, a_i2, ..., a_ik}, a_ij∈I, k≤mand i=1, ..., n.
[0103] In operation (1), an empty root node of a sub-frequent pattern tree may be defined, and the minimum support of the sub-frequent pattern tree (e.g., expected threshold, etc.) may be defined; in operation (2), all items may be counted, but the exemplary embodiment is not limited thereto. The count C of the item (e.g., frequency of occurrence, etc.) is the number of item sets of a contained in D: In operation (3), the infrequently accessed items whose support is less than the minimum support can be discarded to obtain a frequently accessed item set (which can be arranged in descending order).
[0104] Figure 8 A schematic diagram illustrating a frequently accessed item set generated in at least one example embodiment of the inventive concept.
[0105] Reference Figure 8 The access list shown, for example, for the 6 transaction data containing the access item set, that is, the AID column, the indexes are 001 to 006 but not limited to this. The set of items respectively contained in each transaction data is shown in the Items column, and each item shown in the Items column corresponds to the index (key) of the embedded vector. In the set of items respectively contained in each transaction data shown in the third column, the frequently accessed item set is obtained after discarding the non-frequently accessed items whose support is less than the minimum support, and the frequently accessed item set obtained can be arranged in descending order, but the example embodiment is not limited to this. In other words, the transaction data can be compressed by determining whether the access item is a frequently accessed item or an infrequently accessed item in the data set and discarding the infrequently accessed item.
[0106] In operation (4), starting from an empty set, each item in the sorted frequently accessed item set may be read in to generate a sub-frequent pattern tree; if the current item already exists in the current sub-frequent pattern tree, the value of the item may be increased; if the current item does not exist in the tree, a branch may be added, but the example embodiment is not limited thereto.
[0107] In order to facilitate subsequent access correlation mining, a header table structure may be added to the sub-frequent pattern tree, but the example embodiments are not limited thereto.
[0108] Fig. 9 A schematic diagram showing the structure of a header table in at least one example embodiment of the inventive concept.
[0109] Reference Fig. 9 , a header table, i.e., a table of titles of frequent items, each entry of which may include three fields: for example, (1) item name (Items): the same as the node of the sub-frequent pattern tree, the node of the sub-frequent pattern tree is the node item; (2) support count (support count): the number of transactions represented by the partial path to the node; (3) head of node link (Head of nodeLink): a pointer to the first node of the item name contained in the sub-frequent pattern tree, etc., but the example embodiments are not limited to this.
[0110] According to at least one exemplary embodiment of the inventive concept, a global frequent pattern tree can be obtained in the following manner: the sub-frequent pattern trees can be merged into a global frequent pattern tree. For example, the attributes and links in the title table corresponding to each sub-frequent pattern tree can be merged; the "attribute" here can include the project name and / or support count in the title table, etc., but is not limited to this. The project names in each sub-frequent pattern tree can be merged, and the count of the project name and the node link header are updated to merge the sub-frequent pattern trees into a global frequent pattern tree. Therefore, a global frequent pattern tree, a global header table corresponding to the global frequent pattern tree, and / or a global node link list consisting of the global node connection header corresponding to the global frequent pattern tree can be obtained, but the exemplary embodiment is not limited to this.
[0111] According to at least one example embodiment of the inventive concept, a conditional frequent pattern tree can be generated in the following manner: for example, mining can be started from the bottom item of the global header table in sequence; for each item in the global header table corresponding to the global frequent pattern tree, its corresponding conditional frequent pattern library can be found. The conditional frequent pattern library is a sub-frequent pattern tree corresponding to the leaf node to be mined (for example, the current node item). After obtaining this sub-frequent pattern tree, the count of each node in the sub-frequent pattern tree can be set to the count of the leaf node (for example, find the prefix path of the current node, and generate the conditional pattern base (Conditional Pattern Base) of the current node based on each prefix path of the current node), and the nodes with counts lower than the minimum support can be deleted, but the example embodiment is not limited to this. From the conditional frequent pattern library, a set of items with high access relevance, that is, a high access relevance item set, can be recursively mined, but it is not limited to this.
[0112] Fig.10 A schematic diagram illustrating a global frequent pattern tree in at least one example embodiment of the inventive concept.
[0113] Fig.11 A schematic diagram of a process of generating a conditional frequent pattern library corresponding to node I4 in at least one example embodiment of the inventive concept is shown.
[0114] According to at least one exemplary embodiment of the inventive concept, a high-access related item set can be mined in the following manner: For example, if the lowest node item of the global frequent pattern tree is I5, I5 is not considered because it does not reach the minimum support count of 3, so it can be deleted to obtain the following: Fig.10 The global frequent pattern tree shown in FIG. 1 , but the exemplary embodiment is not limited thereto. If the next lower node is I4, assuming that the sub-frequent pattern tree corresponding to node I4 is as follows: Fig.11 As shown, I4 appears in 4 branches, {I2, I1, I3, I4: 1}, {I2, I1, I4: 1}, {I2, I3, I4: 1}, {I4: 1}. When I4 is used as a suffix, the prefix path will be {I2, I1, I3: 1}, {I2, I1: 1}, {I2, I3: 1}, and these prefix paths form the conditional pattern base of node I4; and there are nodes in two of the branches that do not meet the minimum support condition of 3, so it can be considered that I4 only appears in 2 branches, for example, {I2, I1, I3, I4: 1}, {I2, I1, I4: 1}. Therefore, if I4 is used as a suffix, the prefix path will be {I2, I1, I3: 1}, {I2, I1: 1}, which will then form the conditional pattern base of node I4, but the example embodiment is not limited to this.
[0115] The conditional frequent pattern library is regarded as a transaction database, and a conditional frequent pattern tree can be constructed and / or generated based on it. Therefore, the conditional frequent pattern library of node I4 will contain {I2:2, I1:2}, and I3 is not considered because it does not meet the minimum support number. Therefore, the set of all high access correlation items corresponding to node I4, that is, the high access correlation item set is: {I2, I4: 2}, {I1, I4: 2}, {I2, I1, I4: 2}, but the example embodiment is not limited to this.
[0116] For I3, the prefix path is {I2, I1:4}, {I2:1}, which will generate a conditional frequent pattern tree of 2 nodes: {I2:5, I1:4}, and can generate corresponding high-access related item sets: {i2, i3:4}, {i1, i3:3}, {i2, i1, i3:3}, but the example embodiment is not limited to this.
[0117] For I1, the prefix path is {I2:4}, which will generate a single-node conditional frequent pattern tree: {I2:4}, and can generate a corresponding high-access related item set: {I2, I1:4}, but the example embodiment is not limited thereto.
[0118] That is, the above process can obtain the results shown in the following table, but is not limited thereto:
[0119]
[0120] According to at least one example embodiment of the inventive concept, in order to reduce or even eliminate the access error rate, the embedding vectors corresponding to the items with high access relevance can be stored in the same CXL device, and multiple CXLs can work in parallel, but the example embodiment is not limited to this. The storage of the embedding vectors corresponding to the items with high access relevance can meet the uniformity and / or access relevance: (1) uniformity: the computing tasks related to the embedding table can be evenly distributed among multiple PNM devices to achieve and / or ensure that all PNMs are in working state; and / or (2) access relevance: the items with high access relevance can be mined by the multi-support parallel optimization FP-Growth algorithm (MS-PFP-Growth), but the example embodiment is not limited to this. The collection of these items with high access relevance can be stored in the form of a conditional frequent pattern tree in the adjacent position of the CXL to facilitate access and calculation of the recommendation algorithm, etc.
[0121] Fig.12 A schematic diagram illustrating storing highly accessed related item sets in the form of a conditional frequent pattern tree in at least one example embodiment of the inventive concept.
[0122] Reference Fig.12 According to at least one example embodiment of the inventive concept, for the access items in the diagram, respective conditional frequent pattern trees can be mined, and therefore, the embedding vectors corresponding to the items contained in the conditional frequent pattern trees corresponding to the items can be centrally stored in adjacent locations of the CXL, and multiple CXLs can be used in parallel to expand the storage capacity, but the example embodiments are not limited thereto.
[0123] According to at least one example embodiment of the inventive concept, when the storage method of the embedded table of at least one example embodiment of the inventive concept is applied to a recommendation system, the performance of the recommendation system can be improved: the MS-PFP-Growth algorithm can be used to mine items with high access correlation, and / or the embedding vectors corresponding to these items can be unloaded to the memory of the CXL-PNM device, which can reduce the access error rate of the embedded table and / or improve the performance of the recommendation system, etc.
[0124] According to at least one example embodiment of the inventive concept, the total cost of ownership (TCO) can be reduced: the cost of CXL-PNM devices is lower than that of GPUs, and, compared with a multi-GPU solution, a solution using CXL-PNM devices does not have a communication bottleneck, but example embodiments are not limited thereto.
[0125] According to at least one example embodiment of the inventive concept, it is possible to implement and / or ensure that the recommendation system runs efficiently after memory expansion: the embedded table offloading solution using the CXL-PNM device can reduce the data interaction between the CPU and memory data, and the problem of limited embedded table storage in the recommendation system can be improved and / or solved by expanding multiple CXL-PNM devices.
[0126] Fig.13 A block diagram illustrating a storage device embedded with a table in at least one embodiment of the inventive concept.
[0127] Reference Fig.13 At least one example embodiment of the inventive concept also provides a storage device 1300 configured to store an embedded table. The storage device 1300 may include but is not limited to an acquisition unit 1301 and / or a storage unit 1302, etc.
[0128] The acquisition unit 1301 (eg, an acquisition circuit, a processing circuit, an acquisition device, etc.) may acquire at least one embedding table;
[0129] The storage unit 1302 (e.g., a storage device, etc.) can store the embedding table in the memory of at least one CXL-PNM device so that during the embedding table lookup, a parallel search is performed in the memory through the processing circuit of the at least one CXL-PNM device, wherein the at least one CXL-PNM device includes a memory and / or a computing module, etc., but the example embodiments are not limited thereto.
[0130] According to at least one example embodiment of the inventive concept, the storage unit 1302 may store the embedding vector corresponding to at least one item set in the embedding table whose access relevance is greater than an expected and / or preset threshold in the memory of the same CXL-PNM device in at least one CXL-PNM device, wherein the embedding table may be constructed based on a data set including at least one transaction, and each data item in the item set may be included in a transaction of the data set.
[0131] According to at least one example embodiment of the inventive concept, the storage unit 1302 can obtain a set of items whose access relevance is greater than an expected and / or preset threshold value based on a frequent pattern growth algorithm through parallel calculation of the computing modules of at least one CXL-PNM device; the embedding vector corresponding to the obtained set of items whose access relevance is greater than the expected and / or preset threshold value can be stored in the memory of the same CXL-PNM device, but the example embodiments are not limited to this.
[0132] According to at least one example embodiment of the inventive concept, the storage unit 1302 may evenly distribute the computational task of obtaining item sets in the data set whose access correlation is greater than an expected and / or preset threshold to the computational modules of each CXL-PNM device, so as to complete the computational task through the parallel operation of the computational modules of each CXL-PNM device.
[0133] According to at least one example embodiment of the inventive concept, the storage unit 1302 can construct a feature embedding table corresponding to at least one feature of each dimension of the data set; each feature embedding table can be parallelly calculated by the computing module of at least one CXL-PNM device to perform operations including but not limited to the following: construct a sub-frequent pattern tree corresponding to each feature embedding table, and set a corresponding support threshold for each sub-frequent pattern tree; merge the attributes and links of each sub-frequent pattern tree to obtain a global frequent pattern tree; process the obtained global frequent pattern tree from the bottom node items in sequence, for the current node item, according to the sub-frequent pattern tree corresponding to the current node item in the global frequent pattern tree, construct a conditional frequent pattern library corresponding to the current node item; according to the conditional frequent pattern library corresponding to the current node item and the support threshold of the sub-frequent pattern tree corresponding to the current node item, construct a conditional frequent pattern tree corresponding to the current node item; and / or according to the conditional frequent pattern tree corresponding to the current node item, obtain an item set containing the current node item whose access relevance is greater than an expected and / or preset threshold, but the example embodiments are not limited to this.
[0134] It is to be understood that in the example embodiment of the storage device 1300 of the above-mentioned embedded table, the specific implementation process is roughly the same as the example embodiment of the storage method of the above-mentioned embedded table, which will not be repeated here, but the example embodiment is not limited to this. The storage device 1300 of the embedded table can be configured as a combination of dedicated hardware, hardware and / or dedicated firmware for executing one or more methods described herein, respectively. For example, these devices can correspond to a dedicated integrated circuit, and / or can also correspond to a module in which software and hardware are combined. In addition, one or more functions implemented by these devices can also be uniformly executed by components in physical entity devices (e.g., processors, client devices or servers, etc.).
[0135] Fig.14 A block diagram of an electronic device illustrating at least one example embodiment of the inventive concept.
[0136] Reference Fig.14The electronic device 1400 includes at least one memory 1401 and at least one processor 1402 (e.g., a processing circuit, etc.), wherein a set of special-purpose computer-executable instructions (e.g., computer-readable instructions, etc.) is stored in the at least one memory 1401, and when the special-purpose computer-executable instruction set is executed by the at least one processor 1402, a storage method of an embedded table according to some example embodiments of the inventive concept is performed.
[0137] As an example, the electronic device 1400 may be a PC, a tablet device, a personal digital assistant, a smart phone and / or other device capable of executing the above-mentioned computer-readable instruction set. Here, the electronic device 1400 is not necessarily a single electronic device, but may also be any device or circuit collection capable of executing the above-mentioned computer-readable instructions (and / or instruction sets) individually or in combination. The electronic device 1400 may also be part of an integrated control system and / or a system manager, etc., and / or may be configured as a portable electronic device interconnected with a local and / or remote (e.g., via wired and / or wireless transmission) interface, etc.
[0138] In the electronic device 1400, the processor 1402 (e.g., processing circuit, etc.) may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller and / or a microprocessor, etc. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0139] The processor 1402 may execute computer-readable instructions and / or codes stored in the memory 1401, wherein the memory 1401 may also store data. The computer-readable instructions and / or data may also be sent and / or received through at least one network via at least one network interface device, wherein the network interface device may employ at least one known transmission protocol.
[0140] The memory 1401 may be integrated with the processor 1402, for example, by placing RAM and / or flash memory, etc., within an integrated circuit microprocessor, etc. In addition, the memory 1401 may include a separate device, such as an external disk drive, a storage array, and / or any other storage device that can be used by a database system. The memory 1401 and the processor 1402 may be operatively coupled and / or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor 1402 can read files stored in the memory.
[0141] In addition, the electronic device 1400 may also include a video display (such as a liquid crystal display) and / or a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 1400 may be connected to each other via a bus and / or a network.
[0142] According to at least one example embodiment of the inventive concept, a non-transitory computer-readable storage medium storing computer-readable instructions may also be provided, wherein when the computer-readable instructions are executed by at least one computing device, the at least one computing device is prompted to perform the above-mentioned embedded table storage method.
[0143] Examples of non-transitory computer-readable storage media include read-only memory (ROM), random-access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random-access memory (RAM), dynamic random-access memory (DRAM), static random-access memory (SRAM), flash memory, nonvolatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above non-transitory computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system, so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers. It should be noted that the computer-readable instructions can also be used to perform additional operations in addition to the above operations or to perform more specific processing when performing the above operations. The contents of these additional operations and further processing have been mentioned in the description of the relevant methods, so they will not be repeated here to avoid repetition.
[0144] According to at least one example embodiment of the inventive concept, a system is provided including at least one computing device and at least one storage device storing computer-readable instructions, wherein the computer-readable instructions, when executed by the at least one computing device, cause the at least one computing device to execute the above-mentioned method for storing an embedded table.
[0145] It should be noted that the system according to at least one example embodiment of the inventive concept may completely rely on the execution of computer programs and / or computer-readable instructions to implement corresponding functions, that is, each unit corresponds to each operation in the functional architecture of the computer program, so that the entire system is called through a special software package (e.g., lib library) to implement the corresponding functions.
[0146] On the other hand, when the above-mentioned system is implemented in software, firmware, middleware or microcode, the program code or code segment for performing the corresponding operation can be stored in a computer-readable medium such as a non-transitory storage medium, so that at least one processor or at least one computing device can perform the corresponding operation by reading and running the corresponding program code or code segment.
[0147] According to at least one example embodiment of the inventive concept, the storage device may be integrated with the computing device, for example, RAM and / or flash memory are arranged within an integrated circuit microprocessor, etc. In addition, the storage device may include an independent device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The storage device and the computing device may be operatively coupled and / or may communicate with each other, for example, via an I / O port, a network connection, etc., so that the computing device can read the computer-readable instructions stored in the storage device.
[0148] According to at least one example embodiment of the inventive concept, a computer program product is provided, comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the method for storing an embedded table as described above is implemented.
[0149] According to at least one example embodiment of the inventive concept, a storage method and / or apparatus for an embedded table, an electronic device, a non-temporary storage medium, a system, and a computer program product are provided. A CXL-PNM device can be used to store and calculate the embedded table, so as to improve and / or solve problems such as memory limitation and storage constraints caused by excessively large embedded tables through storage expansion based on the CXL-PNM device. When used in a recommendation system, the performance of the recommendation system can be improved. Moreover, through the near storage calculation of the CXL-PNM device, data transmission between processing circuits such as a CPU and memory can be greatly reduced, the storage resources of the CPU can be released and / or used more efficiently, and the operating burden of the CPU can be reduced, so as to improve the execution efficiency of the recommendation system and / or improve and / or optimize the performance of the recommendation system.
[0150] In addition, the embedding vectors corresponding to the itemsets with high access relevance are centrally stored in the CXL-PNM device, which can improve the query efficiency of the embedding vectors.
[0151] In addition, evenly distributing computing tasks among multiple CXL-PNM devices can improve and / or optimize task execution efficiency.
[0152] In addition, through the multi-supported parallel optimization frequent pattern algorithm, it is possible to efficiently mine item sets with high access relevance in data items based on the storage and / or computing architecture of multiple CXL-PNM devices, so as to centrally improve and / or efficiently store the embedding vectors corresponding to the item sets with high access relevance in multiple CXL-PNM devices.
[0153] The above describes some example embodiments of the inventive concept. It should be understood that the above description is only an example and is not exhaustive. The example embodiments of the inventive concept are not limited to the disclosed example embodiments, and many modifications and / or changes may be implemented as example embodiments without departing from the scope and spirit of the inventive concept as defined by the claims herein.
Claims
1. A storage method for an embedded table, characterized in that: The storage method comprises: Get the embedding table; The embedding table is stored in a memory of at least one CXL-PNM device, the storing step causing the at least one CXL-PNM device to perform a parallel lookup of the embedding table in the memory using processing circuitry of the at least one CXL-PNM device, wherein the at least one CXL-PNM device includes the memory and the processing circuitry.
2. The storage method of the embedded table as claimed in claim 1, characterized in that: Also includes: Generating the embedded table based on a data set including at least one transaction, wherein each data item in the data set is included in a transaction of the data set; The embedding vectors corresponding to the item set in the embedding table are stored in the memory in the at least one CXL-PNM device, each item of the data set having an access relevance greater than a desired threshold.
3. The storage method of the embedded table as claimed in claim 2, characterized in that: Also includes: Based on the frequent pattern growth algorithm, we obtain the item sets whose access relevance is greater than the expected threshold; The embedding vector corresponding to the obtained item set is stored in the memory of the at least one CXL-PNM device.
4. The storage method of the embedded table as claimed in claim 3, characterized in that: in, The at least one CXL-PNM device is a plurality of CXL-PNM devices; The method further comprises: Evenly distribute the computational task of obtaining the item sets in the data set whose access relevance is greater than the expected threshold to the processing circuit of each of the plurality of CXL-PNM devices; The computational tasks are accomplished by the parallel operation of the processing circuitry of each of the plurality of CXL-PNM devices.
5. The storage method of the embedded table as claimed in claim 3, characterized in that: Also includes: For each dimension of the data set, generating a feature embedding table, each feature in the feature embedding table corresponds to a feature of the corresponding dimension; Using the processing circuit of the at least one CXL-PNM device to perform parallel calculations on each of the feature embedding tables, the parallel calculations comprising: Generate a sub-frequent pattern tree corresponding to each of the feature embedding tables, For each of the sub-frequent pattern trees, a corresponding support threshold is set. Merge the attributes and links of each of the sub-frequent pattern trees to obtain a global frequent pattern tree, The obtained global frequent pattern tree is sequentially processed from a bottom layer of the obtained global frequent pattern tree to multiple node items included in the obtained global frequent pattern tree. The sequential processing steps include: For a current node item among the multiple node items, Based on the sub-frequent pattern tree corresponding to the current node item in the global frequent pattern tree, a conditional frequent pattern library corresponding to the current node item is generated, Based on the conditional frequent pattern library corresponding to the current node item and the support threshold of the sub-frequent pattern tree corresponding to the current node item, a conditional frequent pattern tree corresponding to the current node item is generated, Based on the conditional frequent pattern tree corresponding to the current node item, an item set including the current node item and having an access relevance greater than an expected threshold is obtained.
6. A storage device embedded in a table, characterized in that The storage device comprises: The first processing circuit is configured to: Get the embedding table; The embedding table is stored in a memory of at least one CXL-PNM device, the storing step causing the at least one CXL-PNM device to perform a parallel lookup of the embedding table in the memory using a second processing circuit of the at least one CXL-PNM device, wherein the at least one CXL-PNM device includes the memory and the second processing circuit.
7. An electronic device, characterized in that: include: at least one processor; at least one memory configured to store computer executable instructions, The at least one processor is configured to execute the computer executable instructions to perform the embedded table storage method of claim 1.
8. A non-transitory computer-readable storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by at least one processor, the at least one processor is prompted to perform the embedded table storage method according to claim 1.
9. A system comprising at least one computing device and at least one storage device storing instructions, characterized in that: in, The at least one computing device is configured to execute computer-readable instructions that cause the at least one computing device to perform the embedded table storage method of claim 1 .
10. A non-transitory computer-readable storage medium storing computer-readable instructions, characterized in that: in, When the computer-readable instructions are executed by at least one processor, the method for storing an embedded table according to claim 1 is implemented.