Dynamic contiguous address memory allocator based on two-level binary tree vectors
Through a two-layer binary tree vector memory allocator, the problem of limited memory access speed of traditional memory allocators in deep learning models is solved, achieving efficient memory management and reducing fragmentation. It is suitable for scenarios such as core switching networks and deep learning.
Patent Information
- Application Number
- CN202411751811.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-12-02
AI Technical Summary
Traditional memory allocators are unable to meet the requirements of efficient memory access in deep learning models, resulting in limited memory access speed and forming a performance bottleneck.
A dynamic continuous address memory allocator based on a two-layer binary tree vector is adopted, including a top-level vector table module, a sub-layer vector table module, a free block positioning module, a first zero coordinate search module and a bit flip module. Memory management is optimized through layering, grouping and resource sharing.
It improves memory allocation efficiency and speed, reduces memory fragmentation, and can efficiently manage large memory spaces. It is suitable for scenarios with high storage requirements such as core switching networks and deep learning.
Smart Images

Figure CN119690659B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of address memory allocation, in particular to a dynamic continuous address memory allocator based on a two-layer binary tree vector. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, especially the breakthrough progress in the field of deep learning, the demand for computing resources, especially memory resources, has increased dramatically. A large number of weight parameters and activation functions are contained in a deep learning model, and these data need to be stored in memory and frequently accessed and updated. The traditional memory allocator cannot meet the efficiency and appears to have a'memory wall' phenomenon, which limits the performance of the operation unit due to the memory access speed and cannot fully exert the computing capacity, forming a performance bottleneck. SUMMARY
[0003] The application provides a dynamic continuous address memory allocator based on a two-layer binary tree vector, aiming to solve the problem of limited memory access speed.
[0004] In order to solve the above technical problems, the technical scheme adopted by the application is as follows: a dynamic continuous address memory allocator based on a two-layer binary tree vector, comprising: a top-level vector table module, a sub-level vector table module, a locating idle block module, a finding first zero coordinate module and a bit flipping module.
[0005] The top-level vector table module comprises a plurality of top-level vector groups, the sub-level vector table module comprises a plurality of sub-level vector groups, the number of top-level vector groups is the same as the number of sub-level vector groups, and the top-level vector groups map the sub-level vector groups to manage the sub-level vectors in the sub-level vector groups.
[0006] The locating idle block module is used for generating an or gate binary tree structure based on the input top-level vector or sub-level vector, and selecting one level vector of the or gate binary tree based on the number of vectors allocated by an external request to output, and locating the idle vector block.
[0007] The finding first zero coordinate module is used for taking the level vector output by the locating idle block module as the input vector of the module, dividing the input vector into a plurality of groups, generating a plurality of and gate binary trees based on the and gate, and finding the first zero vector coordinate.
[0008] The bit flipping module is used for flipping the allocated vector value to 1 and flipping the released vector value to 0, and completing the occupation and release of the vector.
[0009] Further, each group of top-level vector group includes a top-level vector and a free level; the top-level vector describes whether the corresponding sub-level vector group can be used in the request allocation; the free level describes the level of the number of allocable continuous free vectors in the corresponding sub-level vector group, indicating the maximum allocable vector number in the corresponding sub-level vector group.
[0010] Further, the top-level vector table module includes 64 groups of top-level vector groups; each group of top-level vector group includes a 1-bit top-level vector and a 3-bit free level.
[0011] Further, each group of sub-level vector group includes a plurality of sub-level vectors; each sub-level vector is mapped to a physical memory space, and the physical memory space is managed by occupying and releasing the sub-level vectors.
[0012] Further, the sub-level vector table module includes 64 groups of sub-level vector groups, each group of sub-level vector group includes 64 bits, the entire sub-level vector table is composed of 64x64=4096 1-bit vectors, and there are 4k-bit sub-level vectors.
[0013] Further, each node of the AND gate binary tree is composed of two bits, and the two bits respectively represent the result of the AND operation of the child nodes and the direction of the zero in the child nodes.
[0014] Further, the locating free block module, in the output hierarchical vector, a zero indicates that the required number of continuous free vectors can be allocated in the vector block of the root node, thereby locating the free vector block for allocation.
[0015] Further, in the first zero coordinate searching module, searching for the first zero vector coordinate includes: searching for the first zero of each group of AND gate binary tree, and outputting the zero searching result of each group; obtaining the index of the first zero vector of each group from the zero searching result of each group through a priority decoder, and calculating the index of the first zero vector in the sub-level vector group based on the hierarchical depth, thereby obtaining the first zero vector coordinate for allocation.
[0016] Further, in the bit flipping module, the input vector is divided into 8 vector groups, and each group includes 8 vectors.
[0017] Further, in the bit flipping module, the vector group includes a global group, a local group, and a local mask group; in the global group, the 8 vectors corresponding to the bit of 1 are all flipped; in the local group, the 8 vectors corresponding to the bit of 1 are only partially flipped; in the local mask group, 8 bits are used to represent whether the 8 vectors corresponding to the bit of 1 in the local group are flipped.
[0018] The application has the beneficial effects that the memory allocator implements dynamic continuous vector allocation and release based on two-layer binary tree vectors, can expand the number of vectors in each layer to manage larger memory space, and reduces the hardware resource overhead and improves the main frequency through layering, grouping, algorithm optimization, resource sharing and space-time multiplexing, has high memory allocation efficiency and speed, reduces memory fragmentation, and can provide an efficient memory management scheme for scenarios with extremely high storage requirements such as core switching networks and deep learning. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor based on the drawings shown.
[0020] Figure 1 The figure is a whole architecture diagram of the memory allocator of the embodiment of the present application.
[0021] Figure 2 The figure is a 64-vector OR binary tree structure diagram of the embodiment of the present application.
[0022] Figure 3 The figure is a double-value node and zero searching circuit diagram of the embodiment of the present application.
[0023] Figure 4 The figure is a grouping flip schematic diagram of the embodiment of the present application.
[0024] Figure 5 The figure is a system workflow diagram of the embodiment of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0026] It should be noted that the description of "first", "second" and the like in the present application is only for the purpose of description and cannot be understood as indicating or implying the relative importance of the technical features indicated or implying the number of technical features indicated. Therefore, the features defined as "first", "second" can be explicitly or implicitly included at least one of the features. In addition, the technical solutions of various embodiments can be combined with each other, but it must be based on the realization of ordinary skilled in the art, when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, nor within the scope of protection required by the present application.
[0027] As shown in Figure 1 Embodiments of the present application are: a dynamic continuous address memory allocator based on two-layer binary tree vector, comprising: a top vector table module, a sub-layer vector table module, a positioning free block module, a first zero coordinate searching module and a bit flipping module.
[0028] The top vector table module includes a plurality of top vector groups, the sub-layer vector table module includes a plurality of sub-layer vector groups, the number of top vector groups is the same as the number of sub-layer vector groups, and the top vector group maps the sub-layer vector group to manage the sub-layer vector in the sub-layer vector group.
[0029] Each top vector group includes a top vector and a free level, the top vector describes whether the corresponding sub-layer vector group can be used in the request allocation, and the free level describes the level of the number of allocatable continuous free vectors in the corresponding sub-layer vector group, indicating the maximum number of allocatable vectors in the corresponding sub-layer vector group. The top vector table module includes 64 top vector groups; each top vector group includes a 1-bit top vector and a 3-bit free level.
[0030] The top vector table module has a structure as shown in Figure 1 It is mainly composed of 64 top vectors, that is, 64 top vectors (TBV) and free levels (FL). Each TBV is 1 bit, which is used to describe whether the corresponding sub-layer vector group can be used in the request allocation. The value of each TBV is updated according to the requested vector number Size and FL when requesting allocation or releasing vector, and the updated value is 1 if it is not available, otherwise 0 if it is available. Each FL is 3 bits, which is used to describe the level of the number of allocatable continuous free vectors in the corresponding sub-layer vector group, indicating the maximum number of allocatable vectors in the group, the value range is 0 to 7, and 7 indicates that there is no free vector in the corresponding sub-layer vector group. When the value of FL[i] is not 7, the maximum number of continuous allocatable vectors in the corresponding SBV is 2 6-FL[i] .
[0031] The generation method of TBV is:
[0032] 1. If FL[i] is 7, the generated TBV[i] is 1;
[0033] 2. If FL[i] is not 7, Size < 64 and 2 6-FL[i] ≥ Size, TBV[i] is 0, otherwise 1;
[0034] 3. If Size > 64 and FL[i] = 0, TBV[i] is 0, otherwise 1;
[0035] When the TBV vector input to the bit flip module is flipped, the corresponding FL change method is:
[0036] 1. If the number of flipped TBV is 1, FL remains unchanged;
[0037] 2. If the request type is allocation and the number of flipped TBV is multiple, the corresponding FL is changed to 7 except the last one. For example, TBV[4] ~ TBV[7] are flipped from 0 to 1, FL[4]
[0038] ~ FL[6] are set to 7, and FL[7] remains unchanged;
[0039] 3. If the request type is release and the number of flipped TBV is multiple, the corresponding FL is changed to 0 except the last one. For example, TBV[2] ~ TBV[3] are flipped from 1 to 0, FL[2]
[0040] is set to 0, and FL[3] remains unchanged;
[0041] Since the allocated or released vector is continuous in space, if the number of vectors is greater than 64, multiple TBVs need to be flipped, such as TBV[i] ~ TBV[k], which will definitely cause TBV[i] ~ TBV[k-1] to be flipped. By the above method of updating FL, FL[i] ~ FL[k-1] can be directly updated, avoiding the bit flip module to update FL[i] ~ FL[k] one by one, greatly improving the update efficiency.
[0042] For example, after power-on initialization, all TBV and FL are 0, and if 130 vectors are requested to be allocated, i.e. 3 TBVs are involved in this allocation, the finally allocated SBV is: SBV[2][1:0], SBV[1][63:0], and SBV[0][63:0]. By the above method, FL[0] and FL[1] can be changed to 7. Since SBV[2] is occupied, the maximum number of continuous zero vectors that can be allocated in SBV[2] is 32. According to formula (1) described below, FL[2] = 1.
[0043] Each sub-layer vector group includes multiple sub-layer vectors; each sub-layer vector is mapped to a physical memory space, and the physical memory space is managed by occupying and releasing the sub-layer vectors. The sub-layer vector table module includes 64 sub-layer vector groups, each sub-layer vector group includes 64 bits, and the entire sub-layer vector table consists of 64×64=4096 1-bit sub-layer vectors, for a total of 4k-bit sub-layer vectors.
[0044] Sub-layer vector table module: structure is as follows Figure 1 As shown in , it mainly consists of 64 groups of sub-layer vectors, namely SBV[0] to SBV
[63] . Each group of SBV is a 64-bit vector, and each vector is mapped to a physical memory space. The memory allocator has a total of 4096 SBV vectors. By occupying and releasing these sub-layer vectors, it manages 4k physical memory spaces. SBV[i], TBV[i] and FL[i] have a one-to-one correspondence. FL[i] is obtained based on the AND gate binary tree of SBV[i]:
[0045] FL[i]=6-ceil(log2(MACB))#(1)
[0046] Formula (1) can be used to update the corresponding FL with the updated SBV. In formula (1), ceil refers to rounding up, and MACB refers to the maximum number of allocable consecutive zero vectors in SBV[i]. Figure 2 As shown (the shaded squares in the figure indicate that the vector value is 1, i.e., occupied), in the group vector SBV[i], i.e., level 6, vectors [2] to [9] are 0, i.e., idle, and the rest are occupied. From the binary tree generated by the OR gate, we know that vectors [1] to [4] in level 5 are 0, vector [1] in level 4 is 0, and the vectors of all other nodes are 0. In the binary tree of the AND gate, the highest level where the zero vector is located is 4, and the number of vectors under any node of level 4 is 4. It can be obtained that MACB is 4, and then FL[i] = 6-ceil(log2(4)) = 4.
[0047] The idle block positioning module is used to generate an OR gate binary tree structure based on the OR gate for the input top-level vector or sub-layer vector, and select a level vector of the OR gate binary tree for output based on the number of vectors allocated by external request, and locate the idle vector block.
[0048] Wherein, in the output hierarchical vector of the idle block positioning module, zero indicates that the vector block of its root node can allocate the number of consecutive idle vectors required this time, thereby locating the idle vector block for allocation.
[0049] Positioning free block module: structure as Figure 1As shown in the figure, based on the 64 vectors input into the module, an or gate is generated according to a binary tree structure, combined with the number of vectors requested to be located LocSize, the module outputs the level L of the binary tree and the vectors of the level L. The calculation method of LocSize and level L is as follows (where n is an integer):
[0050]
[0051]
[0052] L = 6 - ceil (log2 (LocSize)) # (4)
[0053] Based on the allocation idea of the binary tree, the number of bottom vectors under each node is a power of 2, so the value of LocSize is equivalent to quantizing LSize by a power of 2 and taking the value upward. For example, when LSize = 12, the node vector of level 2 has 16 bottom vectors, which can just meet the allocation for LSize = 12, so LocSize is a kind of quantization method for LSize.
[0054] From formulas (3) and (4), it is known that when LSize is 1, the output vector of level L = 6, i.e. the bottom vector, is 64; when LSize is 2, the output vector of level L = 5 is 32; when LSize is 3 or 4, the output vector of level L = 4 is 16; when LSize is 5-8, the output vector of level L = 3 is 8; when LSize is 9-16, the output vector of level L = 2 is 4; when LSize is 17-32, the output vector of level L = 1 is 5; when LSize is 33 to 64, the output vector of level L = 0 is 1. Based on formula (4), the multi-way selector selects the vector of level L as the output vector, and since the number of vectors of each level is different, high-bit zero padding is needed for levels 0-5 to meet the number of output vectors is 64. Thus, level L and the vector of the level can be obtained according to Size and the input vector source, and the bottom vector under any node vector of level L is a block of continuous vectors, so the free block can be located.
[0055] For example, in Figure 2 , the vector [1] in level 2, i.e. the second vector from the left in level 2, has the bottom vector
[16] -
[31] under it, i.e. a block of continuous regions of bottom vectors can be located from the level vector.
[0056] The first zero coordinate finding module is used to take the level vector output by the free block locating module as the input vector of the module, divide the input vector into multiple groups, and generate multiple and gate binary trees based on the and gate, to find the first zero vector coordinate;
[0057] Each node of the AND gate binary tree consists of two bits, which respectively represent the result of the AND operation of the child nodes and the direction of the zero in the child nodes.
[0058] The first zero vector coordinate searching module includes: searching for the first zero of each group of AND gate binary trees, and outputting the zero searching results of each group; obtaining the index of the first zero vector of the first group from the zero searching results of each group through a priority decoder, and calculating the index of the first zero vector in the child layer vector group based on the hierarchical depth, thereby obtaining the allocated first zero vector coordinate.
[0059] The first zero vector coordinate searching module has a structure as shown in Figure 1 The module is inputted with 64 vectors from the positioning idle module, and the 64 vectors are divided into 4 groups. The first zero of each group of 16 vectors is searched after generating the AND gate binary tree, and the zero searching results of the four groups are outputted and shifted according to the group number and the hierarchical level L of the input vector from the positioning idle block. Finally, the lowest position group result is outputted from the effective results of the shifted groups through a multiplexer. The zero searching circuit of each group is as shown in Figure 3 Each group of 16 vectors is generated into an AND gate binary tree. The left value a of each node (except the lowest layer) is the AND gate result, and the right value b is the zero direction. When the a value of the left child node of the node is 0, b takes 0, otherwise b takes 1. Taking (a (i)(k) , b (i)(k) ) as a node (i is the hierarchical height of the binary tree layer where the node is located, and k is the left coordinate), its child nodes are (a (i-1)(2k) , b (i-1)(2k) ) and (a (i-1)(2k+1) , b (i-1)(2k+1) ), and (a (i)(k) , b (i)(k) ) is expressed as:
[0060] a (i)(k) = a (i-1)(2k) AND a (i-1)(2k+1) #(5)
[0061] b (i)(k) = a (i-1)(2k) #(6)
[0062] From formulas (5) and (6), the two values (a, b) of each binary tree node can be obtained. Since b represents the direction of the zero in the child nodes, the bs of each hierarchical level are used as the input signals of the multiplexer, and the selection results of the higher hierarchical levels are used as the selection signals of the multiplexer, so that the coordinates of the first zero vector in the 16 vectors in the group can be calculated; and the a value of the highest hierarchical level is a (0)(0)The AND gate result of the group of 16 vectors represents whether there is a zero vector in the group of vectors, that is, whether the coordinates of the first zero vector output by the zero detection circuit of the group are valid. The 4 groups of zero detection circuits in the module respectively search for the first zero in each group of 64 vectors, which can reduce the length of the logic chain in the combinational logic circuit and achieve a higher circuit frequency.
[0063] The 4 groups of zero detection circuits respectively output 4 first zero coordinates to the splicing shift circuit, the splicing circuit splices the 4 groups of first zero coordinates with the group number, and the shift circuit is shifted based on the level L, and the representation method is as follows:
[0064] Splicing result = {group number of the group, first zero coordinates of the group number}; shift result = splicing result << (6-L). Where << represents left shift symbol, and left shift (6-L) is equivalent to multiplication by 2 6-L .
[0065] The 4 groups of shift results are used as inputs of the priority encoder, and the control signal of the priority encoder is a (0)(0) of each group. The group number 0 has the highest priority, followed by the group number 1, and the first zero coordinate value found by the module is output in the order of priority. If the source vector of the idle block positioning is TBV, the first zero coordinate value output by the idle block positioning after being processed by the module will be stored in the high address register (HAddr); if the source vector of the idle block positioning is SBV, the first zero coordinate value output by the idle block positioning after being processed by the module will be stored in the low address register (LAddr).
[0066] For example, when the group number 0 (whose vector coordinates are located in the [0, 15] interval) a (0)(0) is 1, that is, the output first zero coordinates are invalid, the first zero coordinates of the 16 vectors of the group number 1 (whose vector coordinates are located in the [16, 31] interval) are 3 and a (0)(0) is 0, and L = 6, then the splicing result output by the module is {2' b01, 4' b0011} = 6' b010011 = 6' d19, and the shift result is equal to the splicing result, that is, the coordinates of the first zero element obtained by splicing the group number 1 and the coordinates 3 are 19.
[0067] Therefore, the first zero coordinates are output by grouping, splicing and shifting. On the one hand, the grouping idea can reduce the length of the combinational logic chain and improve the system frequency. On the other hand, the priority mechanism of the first zero can reduce the fragmentation of the memory and improve the memory utilization.
[0068] The bit flip module is used to flip the allocated vector value to 1 and the released vector value to 0, and complete the occupation and release of the vector.
[0069] In the bit flip module, the input vectors are divided into 8 vector groups, each group having 8 vectors. In the bit flip module, the vector groups include a global group, a local group, and a local mask group; in the global group, the 8 vectors corresponding to a bit of 1 are all flipped; in the local group, the 8 vectors corresponding to a bit of 1 are only partially flipped; and in the local mask group, 8 bits are used to represent whether the 8 vectors corresponding to a bit of 1 in the local group are flipped.
[0070] The bit flip module takes vectors from a top-level vector table or a sub-level vector table, flips the vectors according to Size and the first zero coordinate (i.e., address Addr), flips the vector value to 1 if the request type is allocation, and flips the vector value to 0 if the request type is release. The bit flip module is internally implemented by a combination of logic and outputs the updated vectors. In the bit flip module, the input 64 vectors are divided into 8 groups, each group having 8 vectors, such as Figure 1 a global group and a local group. In the global group, the 8 vectors corresponding to a bit of 1 are all flipped; and in the local group, the 8 vectors corresponding to a bit of 1 are only partially flipped. Due to the characteristics of binary tree allocation, there can be multiple consecutive bits of 1 in the global group, but there is at most one bit of 1 in the local group. Whether the 8 vectors corresponding to a bit of 1 in the local group are flipped is represented by 8 bits of a local mask, as shown in Figure 4 The algorithms for calculating the global group GlobalGroup, the local group LocalGroup, and the local mask LocalMask are as follows:
[0071]
[0072] LocalMask = 2 (CSize%8)+(CAddr%8) -2 CAddr%8 #(11)
[0073] In formulas (7), (8), Size[12:0] is the number of vectors of an external request, the value of which is 1-4096, and the range of the obtained CSize value is 1-64. HAddr and LAddr are high and low addresses output by the zero search first coordinate module, and the value range is 0-63. According to different input vectors from different sources, the value of CSize and CAddr is different. CSize represents the number of vectors that need to be flipped in the 64 vectors input into the module, and CAddr represents the first flipped vector address (index).
[0074] In formulas (9), (10), (11), the floor function is a down rounding. In the circuit, the down rounding operation of division by 8 can be realized by bit truncation, taking bits 3 and above, i.e., [MSB:3], MSB representing the most significant bit; and nThe operation can be achieved by shifting 1 to the left by n bits; the operation of taking remainder 8 can be achieved by bit truncation, i.e. only taking bits [2:0]. Based on addition and subtraction, shifting and truncation, formula (9), (10) and (11) can be achieved to obtain the global group GlobalGroup, the local group LocalGroup and the local mask LocalMask, and then the 64 vectors input into the module are determined and updated by the following methods:
[0075] 1. If the vector is located in the 8 vectors corresponding to bit 1 in the global group GlobalGroup, the request type is allocation, then the vector is flipped to 1, and the request type is release, then the vector is flipped to 0;
[0076] 2. If the vector is located in the 8 vectors corresponding to bit 1 in the local group LocalGroup, and the position of the vector in the 8 vectors is 1 in the local mask LocalMask, the request type is allocation, then the vector is flipped to 1, and the request type is release, then the vector is flipped to 0;
[0077] 3. If the above 1 and 2 are not true, the vector remains unchanged.
[0078] For example, when CSize = 18 and CAddr = 16, GlobalGroup = 12, corresponding to binary 8' b00001100, i.e. all vectors in groups [2] and [3] are flipped, and the input vectors
[16] ~
[31] are flipped; LocalGroup = 16, corresponding to binary 8' b00010000, i.e. part of the vectors in group [4] are flipped, and LocalMask = 3, corresponding to binary 8' b00000011, then the vectors [0] [1] of group [4], i.e. the input vectors
[32]
[33] are flipped. As shown in Figure 4 , the input vectors
[16] ~
[33] can be flipped to achieve the function of the module.
[0079] Thus, based on the algorithm more suitable for digital circuit implementation, the 64 vectors input into the module are grouped and flipped, which reduces the consumption of circuit resources and reduces the delay of combinational logic. In the memory allocator, by means of resource sharing and different space-time multiplexing, a bit flip module is used to flip and update the top-level vectors and the sub-level vectors with less circuit resource consumption.
[0080] The working process of each module in the memory allocator is shown in Figure 5 , and the specific steps of the system working process are as follows:
[0081] 1. The system is powered on and initialized to zero, TBV, FL and SBV;
[0082] 2. Wait for external request signal;
[0083] 3. If the request type is allocation and the number of request vectors Size is not more than 64 (only one TBV and its corresponding 64 SBV vectors are needed):
[0084] 3.1 In the top-level vector table module, update the TBV based on the number of request vectors Size and the value of FL, and output the TBV to the locate free block module;
[0085] 3.2 In the locate free block, generate a binary tree based on the TBV by OR gate. Since Size is not more than 64, only the 64 vectors at the bottom level of the binary tree are output to the find first zero coordinate module;
[0086] 3.3 In the find first zero coordinate module, if no zero vector can be found in each group, the memory allocator reports allocation failure to the outside and returns to step 2 (wait for external request signal). If a zero vector is found, output the lowest bit zero coordinate to the high address register and the bit flip module;
[0087] 3.4 In the bit flip module, according to the request type being allocation, flip some vector of the input TBV to 1, and in the top-level vector table module, set the corresponding FL of the flipped vector to 7;
[0088] 3.5 At the same time as step 3.4, in the sub-level vector table, select SBV[HAddr] according to the first zero vector coordinate HAddr output in step 3.3 and output the SBV[HAddr] to the locate free block;
[0089] 3.6 In the locate free block, generate a binary tree based on the SBV by OR gate, and select some level vectors of the binary tree to output 64 vectors after high-bit zero padding according to formulas (2), (3), and (4) to the find first zero coordinate module;
[0090] 3.7 In the find first zero coordinate module, find the coordinate of the first zero vector in each group and output the lowest bit zero coordinate to the low address register LAddr and the bit flip module;
[0091] 3.8 In the bit flip module, according to the request type being allocation, flip some vector of the input SBV to 1, and update the flipped SBV to the sub-level vector table and the locate free block;
[0092] 3.9 In the locate free block, generate a binary tree based on the updated SBV by OR gate, calculate the FL value corresponding to the SBV based on the binary tree, and update the calculated FL value to the top-level vector table. Complete this allocation, output the address as {HAddr, LAddr} (high address register concatenating low address register), and return to step 2 (wait for external request signal);
[0093] 4. If the request type is allocation and the number of request vectors Size is more than 64 (multiple TBVs and corresponding SBVs are needed, as Figure 5 the short dashed branch) :
[0094] 4.1 In the top-level vector table module, update the TBV based on the number of request vectors Size and the value of FL, and output the TBV to the locate free block module;
[0095] 4.2 In the locate free block module, generate a binary tree based on the TBV. Since Size is more than 64, select a certain layer of vectors of the binary tree according to formulas (2), (3), and (4), and output the 64 vectors after high-bit zero padding to the find first zero coordinate module;
[0096] 4.3 Synchronize with step 3.3;
[0097] 4.4 In the bit flip module, according to the request type being allocation, flip the multiple vectors of TBV input into the module to 1, set the corresponding FL of the TBV vectors that are flipped to 7 in the top-level vector table module, set all the corresponding SBVs of the TBV vectors that are flipped to 1 in the sub-level vector table module except for the last one, and output the corresponding SBV of the last one of the TBV vectors that are flipped to the bit flip module for processing;
[0098] 4.5 In the bit flip module, according to the request type being allocation, flip the partial vectors of SBV input into the module to 1, and update the flipped SBV to the sub-level vector table and send it into the locate free block;
[0099] 4.6 Same as 3.9;
[0100] 5. If the request type is release (as Figure 5 the long and short dashed branch) :
[0101] 5.1 In the top-level vector table, update the TBV based on the number of request vectors Size and the value of FL, and output the TBV to the bit flip module;
[0102] 5.2 In the bit flip module, according to the request type being release, flip the single or multiple vectors of TBV to 0, set the corresponding FL of the TBV vectors that are flipped to 0 in the top-level vector table module, set all the corresponding SBVs of the TBV vectors that are flipped to 0 in the sub-level vector table module except for the last one, and output the corresponding SBV of the last one of the TBV vectors that are flipped to the bit flip module for processing;
[0103] 5.3 In the bit flip module, according to the request type being release, flip the related vectors of SBV input into the module to 0, and update the flipped SBV to the sub-level vector table and send it into the locate free block.
[0104] 5.4 In the positioning of the idle block according to the updated SBV value by or gate generation binary tree, based on the binary tree calculation the SBV corresponding FL value, and update the calculated FL value to the top vector table. Complete this release, and back to the flow step 2 (waiting for external request signal).
[0105] Based on the above workflow, when the number of vectors requested by the external allocation does not exceed 64, 8 cycles can be completed; when the number of vectors requested by the external allocation exceeds 64, 6 cycles can be completed; when the external request releases the vector, 4 cycles can be completed.
[0106] In summary, the memory allocator implements dynamic continuous address (vector) allocation and release based on two-layer binary tree vector, which can expand the number of vectors in each layer to manage larger memory space, and reduce the overhead of hardware resources and improve the main frequency through layering, grouping, algorithm optimization, resource sharing and time-space multiplexing, etc. It has high memory allocation efficiency and speed and reduces memory fragmentation, providing an efficient memory management solution for scenarios with extremely high storage requirements such as core switching networks, deep learning, etc.
[0107] The above only describes the embodiments of the present application, and does not limit the patent scope of the present application, any equivalent structure or equivalent flow transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A dynamic continuous address memory allocator based on a two-layer binary tree vector, characterized in that: include: Top-level vector table module, sub-level vector table module, free block location module, first zero coordinate search module, and bit flip module; The top-level vector table module includes a plurality of top-level vector groups, and the sub-layer vector table module includes a plurality of sub-layer vector groups. The number of the top-level vector groups is the same as the number of the sub-layer vector groups. The top-level vector groups are mapped to the sub-layer vector groups to manage the sub-layer vectors in the sub-layer vector groups. The idle block positioning module is used to generate an OR gate binary tree structure based on an OR gate for the input top-level vector or sub-layer vector, and select a level vector of the OR gate binary tree for output based on the number of vectors allocated by the external request, and locate it to an idle vector block; The first zero coordinate search module is used to use the hierarchical vector output by the free block positioning module as the input vector of the module, divide the input vector into multiple groups and generate multiple groups of AND gate binary trees based on AND gates to find the first zero vector coordinates; The bit flip module is used to flip the allocated vector value to 1 and flip the released vector value to 0, thereby completing the occupation and release of the vector.
2. The dynamic continuous address memory allocator based on a two-layer binary tree vector as claimed in claim 1, characterized in that: Each top-level vector group includes a top-level vector and an idle level; the top-level vector describes whether the corresponding sub-layer vector group can be used in the requested allocation; the idle level describes the level at which the number of allocable consecutive idle vectors in the corresponding sub-layer vector group is located, indicating the maximum number of allocable vectors in the corresponding sub-layer vector group.
3. The dynamic continuous address memory allocator based on a two-layer binary tree vector as claimed in claim 2, characterized in that: The top-level vector table module includes 64 top-level vector groups; each top-level vector group includes a 1-bit top-level vector and a 3-bit idle level.
4. The dynamic continuous address memory allocator based on a two-layer binary tree vector as claimed in claim 1, characterized in that: Each sub-layer vector group includes multiple sub-layer vectors; each sub-layer vector is mapped to a physical memory space, and the physical memory space is managed by occupying and releasing the sub-layer vectors.
5. The dynamic continuous address memory allocator based on two-layer binary tree vectors according to claim 4, characterized in that: The sub-layer vector table module includes 64 groups of sub-layer vector groups, each group of sub-layer vector groups includes 64 bits, and the entire sub-layer vector table is composed of 64×64=4096 1-bit sub-layer vectors, a total of 4k-bit sub-layer vectors.
6. The dynamic continuous address memory allocator based on two-layer binary tree vectors according to claim 1, characterized in that: Each node of the AND gate binary tree consists of two bits, and the two bits respectively represent the result of ANDing the child nodes and the direction of the zero in the child nodes.
7. The dynamic continuous address memory allocator based on a two-layer binary tree vector as claimed in claim 1, characterized in that: The idle block positioning module outputs a hierarchical vector in which zero indicates that the vector block of its root node can allocate the number of consecutive idle vectors required this time, thereby locating an idle vector block for allocation.
8. The dynamic continuous address memory allocator based on a two-layer binary tree vector as claimed in claim 1, characterized in that: In the module for searching for the first zero coordinates, searching for the first zero vector coordinates includes: searching for the first zero in each group of AND gate binary trees and outputting the zero search results of each group; obtaining the subscript of the first group of first zero vectors from each group of zero search results through a decoder with priority, and calculating the subscript of the first zero vector in the sub-layer vector group based on the hierarchical depth, thereby obtaining the allocated first zero vector coordinates.
9. The dynamic continuous address memory allocator based on a two-layer binary tree vector as claimed in claim 1, characterized in that: In the bit flip module, the input vector is divided into 8 vector groups, each group contains 8 vectors.
10. The dynamic continuous address memory allocator based on two-layer binary tree vectors according to claim 9, characterized in that: In the bit flip module, the vector group includes a global group, a local group, and a local mask group; in the global group, all 8 vectors corresponding to bits that are 1 are flipped; in the local group, only some of the 8 vectors corresponding to bits that are 1 are flipped; in the local mask group, 8 bits are used to indicate whether the 8 vectors corresponding to bits that are 1 in the local group are flipped.
Citation Information
Patent Citations
Memory allocation method for adding and deleting nodes in binary tree
CN112328389A
Memory management method and system based on complete binary tree
CN112596908A