CXL-based optimization tensor transmission method and device, and storage medium
By mounting the consistency cache area on the AI accelerator side and optimizing tensor transmission with CXL, the problem of low computing and communication efficiency in deep learning model training is solved, achieving more efficient training efficiency and cost-reducing effect.
Patent Information
- Application Number
- CN202510292287.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems with low computing and communication efficiency in deep learning model training, especially in the case of limited bandwidth of heterogeneous frameworks and PCIe ports, resulting in redundant transmission and delay.
By mounting the consistency cache area on the AI accelerator side and using CXL (Compute ExpressLink) to implement mapping, the tensor transfer method is optimized. Specific steps include storing the parameters and gradients between the CPU and the AI accelerator in the consistency cache area, and performing cache line updates and out-of-memory access signal processing when cached Miss.
This method reduces the storage burden on the CPU side, improves the efficiency of deep learning training, reduces the cost, and reduces the CXL communication traffic and improves the communication efficiency by introducing an integrator and decomposer.
Smart Images

Figure CN120144501A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data transmission technologies. For example, it relates to an optimized tensor transmission method, device, and storage medium based on CXL. Background Art
[0002] In recent years, deep learning models have become increasingly large, and most scenarios also have higher requirements for large models. However, training these large models faces a memory capacity barrier. Model parameters and intermediate results are prone to exceeding the memory capacity. Distributed model training technology can split the model state for training on artificial intelligence accelerators or even multiple artificial intelligence accelerators.
[0003] However, the use of heterogeneous frameworks will introduce additional resource overhead, and at the same time, the PCIe port bandwidth is limited, which will bring relatively large latency. Redundant transmission is one of the main problems faced in tensor transmission. First, the data flow of parameters in a heterogeneous platform has strong directionality, and at the same time, the size of the parameters is smaller than the cache line size, so the PCIe bandwidth of a single transmission will be wasted.
[0004] Therefore, there are problems of low computing and communication efficiency in existing technical solutions.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application. Summary of the Invention
[0006] To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a general review, nor is it intended to identify key / important elements or delineate the protection scope of these embodiments. Instead, it serves as a preface to the subsequent detailed description.
[0007] The embodiments of the present disclosure provide an optimized tensor transmission method based on CXL. The method includes: Dividing a part of the global memory of the artificial intelligence accelerator into a coherent cache area corresponding to the CPU host side and implementing mapping through CXL. The coherent cache area is used to store the parameters and gradients transmitted between the CPU and the artificial intelligence accelerator, and the remaining tensors are allocated to the non-coherent memory area corresponding to the artificial intelligence accelerator side; When making a data request to the CPU, first initiate a data request to the CPU cache. If the required cache line is valid, the data access is completed. If the required cache line is invalid, initiate a memory access request to the coherent cache area. When there is still no valid cache line, it is regarded as a CPU cache miss.
[0008] In some embodiments, the method further includes: When a CPU cache miss occurs, the cache line is updated, and an external memory access signal Change_valid is sent to the CPU host; After receiving the Change_Valid signal, the CPU host first transmits the Change_Valid signal to the coherent cache area. After receiving the Change_Valid signal, the state of the coherent cache area changes from the exclusive state E to the invalid state I, and a state change confirmation signal State_Ready is fed back to the CPU host; After receiving the State_Ready signal and determining that the external memory device is available, the CPU host sends an external memory data start read-in signal Data_Fetch to the CPU cache to indicate that the external memory data starts to be read in. At the same time, a read request signal Read_En is sent to the memory controller; After receiving the Data_Fetch signal, the CPU cache indicates that the data starts to be read in from the external memory, and the state changes from I to E.
[0009] In some embodiments, the method further includes: The memory controller sends the data Data to the CPU cache, or the CPU host writes it into the CPU cache, entering the cache update stage; Based on the data Data overwriting the original cache line, the cache line state of the CPU cache is changed from E to the modified state M, and an update completion signal Update_Finish is sent to the CPU host.
[0010] In some embodiments, the method further includes: The CPU host sends an enable signal Flush_En that allows data update to the CPU cache; After receiving the Flush_En signal, the CPU cache sends the cached data to the coherent cache area in the form of CXL packets through the PCIe physical layer and CXL; After the data transmission is completed, the cache line state of the CPU cache changes from M to the shared state S, and the cache line state of the coherent cache area also changes from I to S; When the corresponding cache line is popped from the CPU cache, the corresponding cache line state changes from S to I; After the CPU host completes data verification, a data update completion signal Flush_Done is sent to the coherent cache area to mark the completion of the cache line update operation based on CXL. The cache line state in the coherent cache area changes from S to E.
[0011] In some embodiments, the method further includes performing updates based on an integrator and a decomposer; Among them, the integrator includes a four-bit configuration register, where the highest bit indicates whether the integration operation is activated, and the last three bits store the Active_Byte parameter; the integration operation is to integrate the Active_Byte bits of each parameter to be updated into a CXL data row, and the Header of the CXL data row includes the configuration register data; After the cache line is evicted from the CPU cache or a flush request is generated, the CPU host checks whether the cache line is mapped to the coherent cache area. If not, the corresponding cache line is directly written to the CPU memory. If so, the corresponding cache line is sent to the input cache of the CXL port and then enters the integrator to generate a CXL data packet; On the artificial intelligence accelerator side, the decomposer is responsible for unpacking the CXL data packet and updating the parameter to be updated.
[0012] In some embodiments, unpacking the CXL data packet and updating the parameter to be updated includes: Fetch the configuration register data of the CXL data packet header and confirm the parameter value of Active_Byte; The decomposer determines the position of the parameter to be modified according to the cache line address and the label index of the parameter, clears the Active_Byte bit to 0, and then performs a bitwise operation on the parameter data in the CXL data row and the parameter data in the coherent cache area to complete the data update.
[0013] In some embodiments, the payload of the CXL data row is 32 bytes.
[0014] In some embodiments, the method further includes monitoring whether the CXL data interaction on the PCIe is completed by monitoring the cache line status of the coherent cache area and the CPU cache.
[0015] Embodiments of the present disclosure provide an electronic device, which includes at least one processor; And a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned CXL-based optimized tensor transmission method.
[0016] Embodiments of the present disclosure provide a storage medium storing program instructions, which execute the above-mentioned CXL-based optimized tensor transmission method when running.
[0017] The CXL-based optimized tensor transmission method, device, and storage medium provided by the embodiments of the present disclosure can achieve the following technical effects: The present disclosure reduces the storage burden on the CPU side by mounting a coherent cache area on the artificial intelligence accelerator side. At the same time, based on the updated CXL policy, there is no need to design a large-capacity snooping filter or coherent directory, which improves the deep learning training efficiency and reduces the cost. An integrator and a decomposer are introduced at the CXL layer to reduce the CXL communication traffic, improve the communication efficiency, and also reduce the cost.
[0018] The above general description and the following description are only exemplary and explanatory, and are not used to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and wherein: Figure 1 is a schematic flowchart of an optimized tensor transmission method based on CXL provided by an embodiment of the present disclosure; Figure 2 is a schematic diagram of data transmission between a CPU and an artificial intelligence accelerator through CXL provided by an embodiment of the present disclosure; Figure 3 is a schematic diagram of an optimized CXL policy based on a host update request provided by an embodiment of the present disclosure; Figure 4 is a schematic diagram of a CXL data packet provided by an embodiment of the present disclosure; Figure 5 is a schematic flowchart of the overall update process of a cache line provided by an embodiment of the present disclosure; Figure 6 is a schematic structural diagram of an optimized tensor transmission device based on CXL provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the drawings. The attached drawings are for reference and illustration only, and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, numerous details are provided to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be shown in a simplified manner to simplify the drawings.
[0021] In the embodiments of the present disclosure, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so as to implement the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion.
[0022] Unless otherwise specified, the term "plurality" means two or more.
[0023] In the embodiments of the present disclosure, the character " / " indicates that the front and rear objects have an "or" relationship. For example, A / B means: A or B.
[0024] The term "and / or" is a description of the association relationship of an object, indicating that there can be three relationships. For example, A and / or B means: A or B, or, the three relationships of A and B.
[0025] The term "corresponding" may refer to an association relationship or a binding relationship. A corresponding to B means that there is an association relationship or a binding relationship between A and B.
[0026] To solve the above problems, the present disclosure provides an optimized tensor transmission method, device, and storage medium based on Compute ExpressLink (CXL).
[0027] The following describes the optimized tensor transmission method, device, and storage medium based on CXL provided by the embodiments of the present disclosure with reference to the accompanying drawings.
[0028] Figure 1 It is a schematic flowchart of an optimized tensor transmission method based on CXL provided by the embodiments of the present disclosure.
[0029] Combined with Figure 1 As shown, the optimized tensor transmission method based on CXL may include: S101, divide a part of the global memory of the artificial intelligence accelerator into a coherent cache area corresponding to the CPU host side, and implement mapping through CXL. The coherent cache area is used to store the parameters and gradients transmitted between the CPU and the artificial intelligence accelerator, and the remaining tensors are allocated to the non-coherent memory area corresponding to the artificial intelligence accelerator side; S102, when making a data request to the CPU, first initiate a data request to the CPU cache. If the required cache line is valid, the data access is completed. If the required cache line is invalid, initiate a memory access request to the coherent cache area. When there is still no valid cache line, it is regarded as a CPU cache miss.
[0030] In some embodiments, Figure 1 the method in may further include: when a CPU cache miss occurs, performing cache line update and raising an external memory access signal Change_valid to the CPU host; After receiving the Change_Valid signal, the CPU host first transmits the Change_Valid signal to the coherent cache region. After receiving the Change_Valid signal, the state of the coherent cache region changes from the exclusive state E to the invalid state I, and a state change confirmation signal State_Ready is fed back to the CPU host; After receiving the State_Ready signal and determining that the external memory device is available, the CPU host sends an external memory data start read-in signal Data_Fetch to the CPU cache to indicate that external memory data starts to be read in, and at the same time sends a Read_En read request signal to the memory controller; After receiving the Data_Fetch signal, the CPU cache indicates that data starts to be read in from the external memory, and the state changes from I to E.
[0031] In some embodiments, Figure 1 the method in may further include: the memory controller sends the data Data to the CPU cache, or the CPU host writes it into the CPU cache, entering the cache update stage; Based on the data Data overwriting the original cache line, the cache line state of the CPU cache is changed from E to the modified state M, and an update completion signal Update_Finish is sent to the CPU host.
[0032] In some embodiments, Figure 1 the method in may further include: the CPU host sends an enable signal Flush_En that allows data update to the CPU cache; After receiving the Flush_En signal, the CPU cache sends the cached data in the form of CXL data packets to the coherent cache region through the Peripheral Component Interconnect Express (PCIe) physical layer and CXL; After completing the data transmission, the cache line state of the CPU cache changes from M to the shared state S, and the cache line state of the coherent cache region also changes from I to S; When the corresponding cache line is popped from the CPU cache, the corresponding cache line state changes from S to I; After the CPU host completes data verification, it sends a data update completion signal Flush_Done to the coherent cache area to mark the completion of the cache line update action based on CXL. The cache line status in the coherent cache area changes from S to E.
[0033] In some embodiments, Figure 1 the method in [[ ]] may further include performing updates based on an integrator and a decomposer; wherein, the integrator includes a four-bit configuration register, the highest bit represents whether the integration operation is activated, and the last three bits store the Active_Byte parameter; the integration operation is to integrate the Active_Byte bits of each parameter to be updated into a CXL data row, and the Header of the CXL data row includes the configuration register data; When a cache line is evicted from the CPU cache or a flush request is generated, the CPU host checks whether the cache line is mapped to the coherent cache area. If not, it directly writes the corresponding cache line to the CPU memory. If so, it sends the corresponding cache line to the input cache of the CXL port and then enters the integrator to generate a CXL data packet; On the artificial intelligence accelerator side, the decomposer is responsible for unpacking the CXL data packet and updating the parameter to be updated.
[0034] In some embodiments, unpacking the CXL data packet and updating the parameter to be updated includes: grabbing the configuration register data of the CXL data packet header to confirm the parameter value of Active_Byte; The decomposer determines the position of the parameter to be modified according to the cache line address and the label index of the parameter, clears the Active_Byte bit to 0, and then performs a bitwise operation on the parameter data in the CXL data row and the parameter data in the coherent cache area to complete the data update.
[0035] In some embodiments, the payload of the above CXL data row can be 32 bytes.
[0036] In some embodiments, Figure 1 the method in [[ ]] may further include monitoring whether the CXL data interaction on the PCIe is completed by monitoring the cache line status of the coherent cache area and the CPU cache.
[0037] Figure 2 is a schematic diagram of data transmission between a CPU and an artificial intelligence accelerator through CXL provided by an embodiment of the present disclosure, Figure 3 is a schematic diagram of an optimized CXL strategy based on a host update request provided by an embodiment of the present disclosure, Figure 4 is a schematic diagram of a CXL data packet provided by an embodiment of the present disclosure, Figure 5It is a schematic diagram of the overall update process of a cache line provided by an embodiment of the present disclosure.
[0038] Combined with Figures 2 to 5 , for Figure 1 the CXL-based optimized tensor transfer method in
[0039] As Figure 2 shown, in the solution of the present disclosure, part of the global memory of the artificial intelligence accelerator is divided into the CXL coherent region of the CPU host-side cache and mapped through CXL. This cache is used to store the parameters and gradients transmitted between the CPU and the artificial intelligence accelerator, and the remaining tensors are allocated to the remaining global memory on the artificial intelligence accelerator side, that is, the traditional non-coherent memory region. During the CPU data request process, a data request is first sent to the CPU cache. If the required cache line is valid, the data access is completed. If the required cache line is invalid, a memory access request is sent to the coherent cache region. When there is still no valid cache line, it is regarded as a cache miss.
[0040] The size of the coherent cache region is not inherently fixed by hardware, but is dynamically divided before each deep learning training and does not change during a single training. Once information such as the deep learning model and the batch of experimental data is determined, the size of the coherent region Coherent_Size will be dynamically configured to ensure that the region is large enough to accommodate all the transfer tensors between the artificial intelligence accelerator and the CPU, and remains unchanged during the deep learning training. In the worst case without any software optimization, the maximum size of Coherent_Size is the sum of the gradients and parameters.
[0041] In addition, the solution of the present disclosure also optimizes the cache update logic. Since the coherent cache region is mounted on the artificial intelligence accelerator side, any cache line change initiated by the artificial intelligence accelerator can complete the cache line update without passing through the PCIe link with high communication cost. For the CPU host side, due to the existence of the coherent cache region, the cache line change initiated within the coherent region does not need to be synchronized to the cache region of the CPU host side, and the CPU can directly send a data request to the coherent region. However, after the CPU cache undergoes a cache line change, since it may affect the training process, it needs to be mapped to the coherent cache region through the CXL policy. Based on this data update characteristic, the solution of the present disclosure proposes an optimized CXL policy based on the host update request, which is only called when the host initiates a cache change request. The specific process is as Figure 3 shown.
[0042] In the default state, the cache line state of the coherence cache region is the exclusive state (E) because the data it stores does not need to be synchronized with the CPU cache. The default state of the CPU cache is (invalid) I, which means that the mapping update to the coherence cache region has been completed.
[0043] When a Miss occurs in the CPU cache and the cache line needs to be updated, an external memory access signal Change_valid is sent to the CPU host. After receiving the Change_valid signal, the host first sends the Change_Valid signal to the coherence cache region. After receiving the signal, the state of the coherence cache region changes from E to I because there is an update request and the current internal storage is invalid, waiting for the arrival of the cached updated data, and a status change confirmation signal State_Ready is sent back to the host. After receiving the State_Ready signal and determining that the external memory device is available, the host sends a Data_Fetch signal to the CPU cache, indicating that the external memory data starts to be read in, and at the same time sends a Read_En read request signal to the memory controller. After receiving the Data_Fetch signal, the CPU cache indicates that the data starts to be read in from the external memory, and the state changes from I to E.
[0044] The data Data is sent from the memory controller to the CPU cache or directly written by the CPU into the CPU cache. After the data transmission is completed, it officially enters the cache update stage. The new data will overwrite the original cache line, completing the update of the CPU cache, and the state of the CPU cache changes from E to M (modified). After the CPU cache completes the cache line update, the Update_Finish signal is pulled high and sent to the CPU host. At this time, since the coherence cache region has completed the state transition from E to I and is waiting to receive the changed data. The CPU host sends an enable signal Flush_En that allows data update to the CPU cache. After receiving the enable signal being pulled high, the CPU cache sends the data to the coherence cache region in the form of CXL packets through the PCIe physical layer and CXL. After the data transmission is completed, at this time, the cache line states in the CPU cache and the coherence cache should be consistent. The state of the CPU cache line changes from M to S (shared), and the cache line state of the coherence cache region also changes from I to S. When the CPU cache evicts this cache line, the state of this cache line changes to I. After the host completes the data verification, a Flush_Done signal is sent to the coherence cache region, indicating that this CXL-based cache line update operation is completed, and the cache line state in the coherence cache region changes from S to E. Thus, the system returns to the default state.
[0045] When the AI accelerator reads a cache line in the coherent cache region, the cache line status remains E. Therefore, by monitoring the cache line status of the coherent cache region and the CPU cache, it is possible to monitor whether the CXL data interaction on the PCIe is completed to ensure the stability of data and parameters during the training process. This solution introduces the CXLCHECK() system function. When this function is called, it checks the cache line status of the coherent cache region and the CPU cache. When the statuses are I and E respectively, it indicates that the CXL system is in the default state and there is no active parameter data interaction, and CXLCHECK() will return 1.
[0046] In a multi-cache system, one of the major obstacles to introducing a large cache is that the size of the snooping filter or the coherence directory will increase as the cache capacity increases. The optimized CXL strategy based on host update requests proposed in this solution does not face this problem. Since the coherent cache region is large enough, the cache line to be updated must exist in the coherent cache region. At the same time, the update behavior is unidirectional, and there will only be cache line updates from the CPU to the AI accelerator. Therefore, there is no need to design a snooping filter in the coherent cache region.
[0047] Taking the FP32 parameter type with a 4-byte length, which is the most commonly used in machine learning, as an example, the cache line length of the AI platform is usually 64 bytes. Therefore, a single cache line can contain at most 16 parameters. However, since in a single update operation, the parameters to be updated are often distributed in different cache lines, this solution also proposes an adapted parameter integration and update method based on an integrator and a decomposer.
[0048] Since the parameter update request from the host occurs at the beginning of each training iteration cycle, parameter integration is activated at the beginning of a single training iteration cycle, and it is separated from the previous parameter integration operation or initialization operation by Aggregation_Step training iteration cycles. Aggregation_Step is a configurable parameter.
[0049] During the deep learning process, continuous training usually only affects the lower significant bits of the parameters, rather than all digits of the parameters. Based on this feature, this solution introduces the Active_Byte parameter for updating significant bits to improve the update efficiency. Active_Byte can be determined by two methods: manually set by the user after the model accuracy is confirmed, or based on the longest bits to be updated of the parameters to be updated after each parameter integration operation is activated. Active_Byte is an integer multiple of 2 and is rounded up.
[0050] The integrator includes a four-bit configuration register. The highest bit indicates whether the integration operation is activated, and the last three bits store the Active_Byte parameter. The integration operation is to integrate the low Active_Byte bits of each parameter to be updated into a CXL data line. The payload of the CXL data line is usually 32 bytes, and the Header of the data line contains the configuration register data, as Figure 4 shown. All integrated CXL data lines, together with the address data of each parameter cache line and the parameter label index, form a CXL data packet.
[0051] Figure 5 Figure
[0052] shows the overall update process of the cache line. When the cache line is evicted from the CPU cache or a flush request is generated, the CPU host checks whether the cache line is mapped to the coherent cache area. If not, the cache line is directly written to the CPU memory. If so, the cache line is sent to the input cache of the CXL port and then enters the integrator for the generation of the CXL data packet. The integrator and the input cache can be implemented using the idle registers and buffers of the CXL link layer without introducing additional complex designs.
[0052] On the side of the AI accelerator, the decomposer is responsible for unpacking the CXL data packet and updating the parameters to be updated. The unpacker first grabs the configuration register data in the CXL data packet header to confirm the parameter value of Active_Byte. The decomposer determines the position of the parameter to be modified according to the cache line address and the label index of the parameter, clears the low Active_Byte bits to 0, and then performs a bitwise operation on the parameter data in the CXL data line and the parameter data in the coherent cache area to complete the data update.
[0053] The solution of the present disclosure is aimed at the heterogeneous multi-core computing platform of CPU + AI accelerator using CXL, focusing on the application scenario of deep learning. By mounting a large-capacity coherent cache area on the side of the AI accelerator, the storage burden on the CPU side is reduced. At the same time, a CXL policy based on updates is designed based on the characteristics of reinforcement learning, without the need to design a large-capacity snooping filter or coherence directory, improving the deep learning training efficiency while reducing the cost. An integrator and a decomposer are introduced at the CXL layer to reduce the CXL communication traffic, improve the communication efficiency, and reduce the cost.
[0054] Combined with Figure 6As shown, an optimized tensor transmission device 600 based on CXL is further provided in an embodiment of the present disclosure, including a processor 604 and a memory 601. Optionally, the system may further include a communication interface 602 and a bus 603. Among them, the processor 604, the communication interface 602, and the memory 601 can complete communication with each other through the bus 603. The communication interface 602 can be used for information transmission. The processor 604 can call the logical instructions in the memory 601 to execute the optimized tensor transmission method based on CXL in the above embodiment.
[0055] In addition, when the logical instructions in the above-mentioned memory 601 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0056] As a computer-readable storage medium, the memory 601 can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present disclosure. The processor 604 executes functional applications and data processing by running the program instructions / modules stored in the memory 601, that is, implements the optimized tensor transmission method based on CXL in the above embodiment.
[0057] The memory 601 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 601 may include high-speed random access memory and may also include non-volatile memory.
[0058] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are set to be an optimized tensor transmission method based on CXL.
[0059] The above computer-readable storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium.
[0060] The technical solution of the embodiments of the present disclosure can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiments of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, including: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or it may also be a transitory storage medium.
[0061] The above description and the drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. The embodiments only represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or substituted for parts and features of other embodiments. As used in the description of the embodiments, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups of these. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, or device including the element. In this document, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts may refer to the description of the method part.
[0062] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0063] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be merely a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. In addition, the functional units in the embodiments of the present disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0064] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a part thereof, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
[0065] The various embodiments of the systems and technologies described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0066] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0067] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0068] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0069] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0070] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0071] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of this disclosure can be achieved, and no limitations are imposed herein.
[0072] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A CXL-based optimized tensor transmission method, characterized in that: The method comprises: Part of the global memory of the AI accelerator is divided into a consistent cache area corresponding to the CPU host side, and mapped through CXL. The consistent cache area is used to store parameters and gradients transmitted between the CPU and the AI accelerator, and the remaining tensors are allocated to the non-consistent memory area corresponding to the AI accelerator side; When making a data request to the CPU, a data request is first initiated to the CPU cache. If the required cache line is valid, the data access is completed. If the required cache line is invalid, a memory access request is initiated to the consistent cache area. When there is still no valid cache line, it is regarded as a CPU cache Miss.
2. The method according to claim 1, characterized in that The method further comprises: When a CPU cache miss occurs, the cache line is updated and the external memory access signal Change_valid is sent to the CPU host; After receiving the Change_Valid signal, the CPU host first transmits the Change_Valid signal to the coherent cache area. After receiving the Change_Valid signal, the coherent cache area changes its state from the exclusive state E to the invalid state I, and feeds back a state change confirmation signal State_Ready to the CPU host. After receiving the State_Ready signal and determining that the external memory device is available, the CPU host sends the external memory data read-in signal Data_Fetch to the CPU cache and sends the Read_En read request signal to the memory controller. After the CPU cache receives the Data_Fetch signal, it indicates that data starts to be read from the external memory, and the state changes from I to E.
3. The method according to claim 2, characterized in that The method further comprises: The memory controller sends the data to the CPU cache, or the CPU host writes the data to the CPU cache, entering the cache update phase; Based on the data Data, the original cache line is overwritten, the cache line state of the CPU cache is changed from E to the modified state M, and an update completion signal Update_Finish is sent to the CPU host.
4. The method according to claim 3, characterized in that The method further comprises: The CPU host sends an enable signal Flush_En to the CPU cache to allow data update; After receiving the Flush_En signal, the CPU cache sends the cached data in the form of CXL packets to the coherent cache area through the PCIe physical layer and CXL; After the data is sent, the cache line state of the CPU cache changes from M to the shared state S, and the cache line state of the coherent cache area also changes from I to S; When the CPU cache pops the corresponding cache line, the state of the corresponding cache line changes from S to I; After completing the data verification, the CPU host sends a data update completion signal Flush_Done to the consistent cache area to mark the completion of the CXL-based cache line update action, and the cache line state in the consistent cache area changes from S to E.
5. The method according to claim 2, characterized in that: The method further includes updating based on the integrator and the decomposer; The integrator includes a four-bit configuration register, the highest bit indicates whether the integration operation is activated, and the last three bits store the Active_Byte parameter; the integration operation is to integrate the Active_Byte bit of each parameter to be updated into a CXL data row, and the Header of the CXL data row includes the configuration register data; When a cache line is evicted from the CPU cache or a refresh request is generated, the CPU host checks whether the cache line is mapped to the coherent cache area. If not, the corresponding cache line is directly written to the CPU memory. If yes, the corresponding cache line is sent to the input cache of the CXL port, and then enters the integrator to generate a CXL data packet; On the AI accelerator side, the decomposer is responsible for unpacking the CXL data packet and updating the parameters to be updated.
6. The method according to claim 5, characterized in that The step of unpacking the CXL data packet and updating the parameters to be updated includes: Capture the configuration register data of the CXL data packet header and confirm the parameter value of Active_Byte; The decomposer determines the location of the parameter to be modified based on the cache line address and the parameter index, and clears the Active_Byte bit to 0. It then performs a bitwise operation on the parameter data in the CXL data line and the parameter data in the consistency cache area to complete the data update.
7. The method according to claim 5, characterized in that The payload of the CXL data line is 32 bytes.
8. The method according to claim 4, characterized in that The method further comprises: By monitoring the cache line status of the coherent cache area and the CPU cache, the completion of the CXL data interaction on the PCIe is monitored.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Data processing system, method, device, medium and program product
CN120492370A