Dynamic memory management method and terminal equipment

By reusing memory blocks that meet storage needs in a large language model and combining the memory increment application mechanism to optimize memory management, the resource consumption problems caused by frequent memory applications and releases are solved, and processing speed and user experience are improved.

CN120371494APending Publication Date: 2025-07-25HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202410251572.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2024-03-05
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

During the processing process, the large language model frequently applies memory and release consumes a large amount of system resources, resulting in slow processing speed and long wait time for users.

Method used

By reusing memory blocks that meet storage needs in a large language model, reducing the number of memory applications and releases, and combining the memory increment application mechanism, memory management is optimized.

Benefits of technology

It improves the processing speed of large language models, reduces user waiting time, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371494A_ABST
    Figure CN120371494A_ABST
Patent Text Reader

Abstract

The invention discloses a memory dynamic management method and terminal equipment, and relates to the field of terminals, the terminal equipment is deployed with a large language model, the large language model comprises a large number of operators, and the first operator is any one of the operators. After the target model receives the first input data; if it is determined that a first memory block applied for the input tensor and the output tensor of the first operator in the historical round exists, and the size of the first memory block meets the requirement of data generated in the processing process of the round, the first memory block is reused to store the data generated in the processing process of the round; and a large number of memory blocks do not need to be released and applied in each round of processing process. The number of times of memory block application and release is reduced, and the processing speed of the large language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202410066241.X and the invention title "A Method for Dynamic Memory Allocation of Large Language Models" filed with the National Intellectual Property Administration on January 16, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0002] This application relates to the terminal field, and in particular, to a method for dynamic memory management and a terminal device. Background Art

[0003] Large language models are a type of natural language processing technology based on neural networks that can learn and predict the rules and patterns of natural language text, thereby achieving the understanding and generation of natural language, and are widely used in fields such as intelligent voice assistants, chatbots, and artificial intelligence painting.

[0004] Large language models generally include a large number of operators. During the process of a large language model processing a conversation, it is necessary to apply for memory for the tensors of each operator respectively to store data.

[0005] Since large language models include a large number of operators with complex connection relationships, applying for memory for each operator of a large language model once consumes a large amount of system resources.

[0006] How to manage the memory application and release process of large language models to avoid a large amount of resource consumption and improve the processing speed of large language models is a problem that needs to be solved. Summary of the Invention

[0007] Embodiments of this application provide a method for dynamic memory management and a terminal device. During the process of managing the memory application and release of large language models, the number of memory applications and releases can be reduced, resource consumption can be lowered, and the processing speed of large language models can be improved.

[0008] To achieve the above objective, the embodiments of this application adopt the following technical solutions:

[0009] In a first aspect, a method for dynamic memory management is provided, and the method includes:

[0010] Obtain the first input data of the target model. According to the first input data of the target model, obtain the first data corresponding to the first operator in the target model. When it is determined that a first memory block has been applied for the first operator of the target model according to the second input data, if the first memory block meets the storage requirements of the first data, then reuse the first memory block to save the first data.

[0011] In this embodiment, by determining whether the first memory block applied for before obtaining the first input data of the target model meets the storage requirement of the first data of the first operator, if it meets the requirement, there is no need to release the first memory block, and there is no need to apply for a new memory block for the first operator, thereby reducing the number of memory applications and releases, reducing resource consumption, and improving the processing speed of the large language model.

[0012] In a possible embodiment, when it is determined that there is a first memory block, if the first memory block fails to meet the storage requirement for storing the first data, the first memory block is released, and then a second memory block is applied for according to the first memory value required to store the first data.

[0013] In this embodiment, when determining whether the first memory block meets the storage requirement of the first data, if the first memory block fails to meet the storage requirement for storing the first data, an action of memory application and release is performed once, so that the newly applied second memory block can meet the storage requirement for storing the first data.

[0014] In a possible embodiment, before obtaining the first input data of the target model, the second input data of the target model is obtained, and according to the number of word segments of the second input data, a first memory block is applied for the input tensor and output tensor of the first operator in the target model.

[0015] In this embodiment, applying for a first memory block for the input tensor and output tensor of the first operator in the target model according to the number of word segments of the second input data can provide an object for judgment when determining whether the first data of the first operator needs to reapply for memory after obtaining the first input data subsequently.

[0016] In a possible embodiment, the second input data of the target model includes the input data generated during the first round of conversation input by the user, or includes the input data corresponding to regarding the conversation input by the user as the first round of conversation. Before obtaining the second input data of the target model, none of the first operators in the target model have applied for the first memory block, or all the first memory blocks have been released.

[0017] In a possible embodiment, the first input data of the target model includes the input data generated during any round of conversation other than the first round of conversation input by the user. Before obtaining the first input data of the target model, the first operators in the target model have already applied for the first memory block.

[0018] In a possible embodiment, if the number of word segments of the first input data is less than the number of word segments of the second input data, it is determined that the size of the first memory block meets the storage requirement of the first data.

[0019] In this embodiment, by determining the quantitative relationship between the number of word segments of the first input data and the number of word segments of the second input data, it is determined whether the first operator in the target model needs to reapply for memory, without having to judge each first operator in the target model one by one, which can reduce resource consumption.

[0020] In a possible embodiment, according to the number of word segments of the first input data, a first memory value required to store the first data of the first operator is determined. If the first memory value is greater than the size of the first memory block, it is determined that the size of the first memory block does not meet the storage requirement for storing the first data.

[0021] In this embodiment, the first memory value includes: the size of the memory corresponding to the dimensions of the input tensor and output tensor of the first operator. When it is determined that the size of the first memory block does not meet the storage requirement for storing the first data, individual judgments are made for each first operator in the target model to avoid the situation where the first memory block cannot completely store the first data.

[0022] In a possible embodiment, after obtaining the first input data of the target model, if there is no memory applied for the first operator, a second memory block is applied according to the input and output tensors of the first operator and the number of word segments of the first input data, and the second memory block is used to store the first data.

[0023] In this embodiment, when there is no memory applied for the first operator, a second memory block is applied for the first operator to achieve on-demand application.

[0024] In a possible embodiment, a first memory value is determined according to the dimensions of the input tensor and output tensor of the first operator, and the size of the second memory block is determined according to the first memory value.

[0025] In a possible embodiment, the size of the second memory block can be the first memory value plus a fixed memory increment value.

[0026] In this embodiment, by adding a fixed memory increment value on the basis of the first memory value of the first operator, a certain margin can be reserved to reduce the number of memory applications and releases.

[0027] In a possible embodiment, the size of the second memory block can also be where x is the first memory value of the first operator, y is the second memory value, and the size of the second memory block is determined according to the second memory value.

[0028] In this embodiment, by adding an increment that varies according to the numerical size of the first memory value on the basis of the first memory value of the first operator, a certain margin can be reserved to reduce the number of memory applications and releases.

[0029] In a possible implementation, the first memory value of the first operator includes the required memory value of the first operator, and the second memory value of the first operator includes the actually applied memory value of the first operator.

[0030] In a possible implementation, determine the input tensor of the target model according to the number of word segments of the first input data; determine the dimensions of the input tensor and output tensor of the first operator according to the dimension of the input tensor of the target model, the data transfer relationship between the first operators of the target model, the parameters and calculation logic within the first operator.

[0031] In this implementation, determine the dimensions of the input tensor and output tensor of the first operator according to the dimension of the input tensor of the target model, the data transfer relationship between the operators of the target model, the parameters and calculation logic within the operator. Provide support for whether the size of the first memory block meets the need to store the first data of the first operator in the subsequent process.

[0032] In a possible implementation, determine the first output of the target model according to the first input data.

[0033] In a possible implementation, after determining the first output of the target model, if the input data of the target model is not received within a preset duration, release the first memory blocks applied for the input tensor and output tensor of each first operator in the target model.

[0034] In this implementation, after not obtaining the next input data of the target model for a long time, release the first memory blocks applied for the input tensor and output tensor of each operator in the target model, which can reduce the memory occupation of the device.

[0035] In a second aspect, a terminal device is provided, including: a processor and a memory; the memory is used to store computer execution instructions, and when the terminal device runs, the processor executes the computer execution instructions stored in the memory, so that the terminal device executes the method described in any item of the first aspect above.

[0036] In a third aspect, a computer-readable storage medium is provided, in which instructions are stored, and when it runs on a computer, the computer can execute the method described in any item of the first aspect above.

[0037] In a fourth aspect, a computer program product containing instructions is provided, and when it runs on a computer, the computer can execute the method described in any item of the first aspect above.

[0038] Among them, the technical effects brought by any design method in the second aspect to the fourth aspect can refer to the technical effects brought by different design methods in the first aspect, which will not be elaborated here. Description of the Drawings

[0039] Figure 1 It is a schematic structural diagram of a model;

[0040] Figure 2 It is a schematic diagram of memory allocation for a model operator;

[0041] Figure 3 It is a schematic diagram of data transfer relationship between operators applicable to the method provided in the embodiment of the present application;

[0042] Figure 4 It is a schematic diagram of incremental memory allocation applicable to the method provided in the embodiment of the present application;

[0043] Figure 5 It is a schematic diagram of memory reuse applicable to the method provided in the embodiment of the present application;

[0044] Figure 6 It is a schematic diagram of memory expansion applicable to the method provided in the embodiment of the present application;

[0045] Figure 7 It is a schematic diagram of memory expansion applicable to the method provided in the embodiment of the present application;

[0046] Figure 8 It is a schematic diagram of the system architecture applicable to the method provided in the embodiment of the present application;

[0047] Figure 9 It is a flowchart of the method applicable to the method provided in the embodiment of the present application;

[0048] Figure 10 It is a schematic diagram of the hardware structure of a terminal device provided in the embodiment of the present application;

[0049] Figure 11 It is a schematic diagram of a chip system provided in the embodiment of the present application. Detailed implementation manners

[0050] In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "the", "above-mentioned", "this" and "such" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one or more than two (including two). The term "and / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist; for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0051] References to "one embodiment" or "some embodiments" etc. described in this specification mean that specific features, structures or characteristics described in connection with that embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way. The term "connection" includes direct connection and indirect connection, unless otherwise stated. "First" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features.

[0052] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0053] A large language model (LLM) is a neural network-based natural language processing technology that can learn and predict the rules and patterns of natural language text, thereby achieving the understanding and generation of natural language, and is widely used in fields such as intelligent voice assistants, chatbots, and artificial intelligence (AI) painting.

[0054] The basic idea of large language models is to regard natural language text as a kind of sequence data, and input this sequence data into a neural network. Through the calculation and transformation of operators in the multiple network layers of the neural network, the corresponding output data is generated. The neural networks in large language models usually adopt network structures such as recurrent neural network (RNN), long short term memory network (LSTM), and gated recurrent unit (GRU) to process the input data. As Figure 1 shown, these neural network structures have multiple network layers, and there are numerous operators in each network layer; the operator is the basic processing unit of each network layer, defining the structure and operation process of the neural network model and realizing the transfer, transformation, and calculation of data. Common operators include convolution operator, pooling operator, activation function operator, etc. The components of an operator include input, calculation logic, parameters, and output; among them, the input is the data source of the operator, which includes one or more tensors. A tensor represents a container for data, and the data loading capacity of the container is represented by the dimension of the tensor. The input tensors of the operator include information such as features, weights, and biases. The calculation logic is the core part of the operator, which defines the mathematical operations, logical operations, transformation operations, etc. of the operator. The calculation logic determines how the operator converts the input data into output data. Parameters can make the operator have learnable attributes. The output is the calculation result of the operator, which also includes one or more tensors. The output tensors of the operator may include the data processed by the calculation logic and can be used as the input of the next operator that has a data transfer relationship with it. As Figure 2 shown, taking a branch of the network structure shown above Figure 1 as an example, the data transfer relationship between operators will be illustrated. In Figure 2In a branch of a network structure, Operator 1, Operator 2, and Operator 3 are respectively connected to Operator 11. The connection lines represent the data transfer paths. It can be seen that Operator 11 has three input tensors from Operator 1, Operator 2, and Operator 3. The dimensions of the input tensors of Operator 11 are related to the dimensions of the output tensors of Operator 1, Operator 2, and Operator 3. Operator 11 is also connected to Operator 111 and Operator 112. The dimensions of the input tensors of Operator 111 and Operator 112 are respectively related to the dimension of the output tensor of Operator 11. In a large language model, each operator corresponds to one or more input and output tensors. Based on the dimensions of the input tensors of the model, the data transfer relationships between the operators in the model, as well as the computational logic, parameters, etc. inside the operators, the dimensions of the input and output tensors of each operator in the model can be obtained.

[0055] With the development of terminal devices, the performance of terminal devices has been greatly improved, which makes it possible to deploy large language models on terminal devices. When using a large language model on a terminal device, the large language model needs to be loaded into the memory of the terminal device. The memory allocator of the terminal device applies for memory blocks of corresponding sizes for each tensor in the memory according to the dimensions of the input and output tensors of each operator in the large language model. When the terminal device does not use the large language model, this part of the memory blocks is released. When the existing memory application and release mechanism applies for and releases memory according to the input and output tensors corresponding to each operator in the large language model, it usually applies for and releases memory in a timely manner on demand to better utilize the limited memory space of the terminal device.

[0056] In an application scenario, a large language model (hereinafter referred to as the model) is used to implement a multi-turn human-computer dialogue function on a terminal device. For each turn of the multi-turn dialogue, the processing process that the user input dialogue goes through includes: data preprocessing, data prediction processing, data postprocessing, etc. Among them, data preprocessing includes converting the user input dialogue into input data that can be recognized by the model. The user input dialogue can be a voice dialogue or a text dialogue. If the user input dialogue is a text dialogue, the data preprocessing process includes performing word segmentation on the user input text dialogue to obtain multiple word segments (tokens) corresponding to the text dialogue, and encoding the word segments using encoding formats such as one-hot, Word2vec, etc. to obtain word vectors. Multiple word vectors constitute the input data. If the user input dialogue is a voice dialogue, then before processing the text dialogue in the above example, the user input voice is converted into text using automatic speech recognition (ASR).

[0057] Data prediction processing includes determining the dimensions of the input and output tensors of each operator in the model (hereinafter referred to as the tensor dimensions of the operator) according to the number of word segments of the input data, and generating a memory application request based on the tensor dimensions of each operator. The memory application request is used to notify the memory allocator of the device to allocate memory blocks of corresponding sizes for the tensor dimensions of each tensor of each operator in the model. The data is stored in the memory blocks, and the operators in the model read the data from the corresponding memory blocks, perform calculations and conversions, and finally obtain the output data.

[0058] Data post-processing includes decoding the word vectors in the output data to obtain word segments, and then combining the word segments into a language text; after obtaining the language text, it can be displayed on the screen of the device by the display driver, and the language text can also be converted into sound for playback using text-to-speech (TTS).

[0059] Since the length and content of the text input in each round of conversation may be different, the number of word segments of the input data input into the model is also different in different rounds of conversation. That is to say, in each round of conversation, it is necessary to determine the tensor dimensions of each operator in the model according to the number of word segments of the input data, and apply for a memory block of the corresponding size from the device's memory according to the tensor dimensions of each operator.

[0060] Exemplarily, one round of conversation refers to a user inputting a piece of edited voice or text. For example, in the interactive page of a human-machine conversation, the user edits text in the input box and then sends the text to the chat box. The model processes the text in the chat box and then generates a reply text to be displayed in the chat box. The user sending the text edited in the input box to the chat box is one round of conversation.

[0061] Such as Figure 3As shown in (a), in the first round of conversation, the input data of the model includes 10 tokens, that is, the number of word segments of the input data is 10. The required memory value corresponding to the dimension of the input tensor of the model is 10 × 4 = 40 (taking single-precision floating-point type as an example). Operator 1 is connected to the input of the model, and the required memory value corresponding to the dimension of tensor11 of Operator 1 is 10 × 4 = 40. Operator 2 is connected to Operator 1, and the required memory value corresponding to the dimension of tensor21 of Operator 2 is 30 × 4 = 120. After obtaining the required memory value corresponding to the dimension of the tensor of the operator, a memory block of the corresponding size is allocated according to the required memory value. For example, a memory block of size 40 is allocated according to the required memory value of 40 for tensor11, and a memory block of size 120 is allocated according to the required memory value of 120 for tensor21. After the first round of conversation ends, the device releases the memory blocks allocated in the first round of conversation.

[0062] As Figure 3 As shown in (b), the input data for the second round of conversation is: 20 tokens, that is, the number of word segments is 20. Similar to the above example, after obtaining the input tensor of the model, the tensor of each operator in the model can be calculated. For example, the required memory value corresponding to the dimension of the input tensor of the model is 20 × 4 = 80, the required memory value of tensor11 of Operator 1 is 80. Operator 1 is connected to Operator 2, and the required memory value corresponding to the dimension of tensor21 of Operator 2 is 60 × 4 = 240. After obtaining the required memory value corresponding to the dimension of the tensor of the operator, a memory block of the corresponding size is allocated according to the required memory value.

[0063] It can be known that a tensor can be regarded as a data matrix. The number of elements (data) in the data matrix is the dimension of this tensor. According to different data types, the memory occupied by each element in the matrix is also different. Generally, a single-precision floating-point data type occupies 4 bytes of memory space for a single element, a double-precision floating-point data type occupies 8 bytes of memory space for a single element, and a half-precision floating-point data type occupies 2 bytes of memory space for a single element. Taking the single-precision floating-point data type as an example, the size of the memory block required for a tensor with a dimension of 60 is 240 bytes. Of course, in actual applications, the dimension of a tensor will be much larger, and the size of the memory block required for a tensor is generally several kilobytes (KB) to dozens of kilobytes. That is to say, there is a corresponding relationship between the data type and the size of the memory applied for. Different data types result in different sizes of memory blocks required for tensors with the same dimension. In the embodiments of the present application, taking the single-precision floating-point data type as an example, a data in the tensor occupies 4 bytes of memory space. The specific values of the dimensions and memory blocks of the tensors involved in the following examples are similar and will not be specifically described again.

[0064] It can be known that Figure 3 The dimension of the tensor of the operator obtained in [reference] is only for some illustrative examples and is not a specific limitation. The dimension of the tensor of the operator is determined according to the input data of the model, the data transfer relationship between the operators in the model, as well as the calculation logic and parameters within the operator. The data transfer relationship between the operators in the model, the calculation logic and parameters within the operator can be known in advance, but the number of word segments of the input data cannot be known in advance. Therefore, the memory size required for the same tensor of the same operator in different rounds of conversations may be different. In addition, each operator may correspond to one or more input and output tensors. When applying for a memory block of the corresponding size according to the dimension of the tensor of each operator, if the operator corresponds to multiple tensors, a corresponding memory block is applied for each tensor respectively. In Figure 3 It is illustrated by taking each operator corresponding to one tensor as an example. For the case where an operator corresponds to multiple tensors, the embodiments of the present application will not be elaborated.

[0065] Since the number of operators in the model is very large, and the memory blocks required for each operator are different according to the number of word segments in each round of conversation, the existing memory application and release mechanism needs to separately apply for and release memory for the tensor of each operator in each round of multiple rounds of conversations, which will consume a large amount of system resources, resulting in a slow processing speed of the model, a longer waiting time for the user to receive a reply after inputting a conversation, and a reduced user experience.

[0066] An embodiment of the present application provides a method for dynamic memory management, which is applied to a terminal device configured with a large language model. A user can conduct multiple rounds of conversations on the terminal device, and the text generated in each round of conversation serves as the input data for the large language model. In each round of conversation, the dimension of the input tensor of the model can be determined according to the number of word segments of the input data; according to the dimension of the input tensor of the model, the data transfer relationship between operators, the calculation logic within the operator, and the parameters, the dimension of the tensor of each operator in the model can be determined. It can be understood that the larger the dimension of the tensor, the larger the memory block for storing data.

[0067] In the method for dynamic memory management provided by the embodiment of the present application, in the first round of conversation, the size of the memory value required for the tensor of each operator is determined respectively according to the dimension of the tensor of each operator, and memory blocks of corresponding sizes are applied for respectively for the size of the memory value required for the tensor of each operator. After each round of conversation ends, the memory blocks are not released first. In the next round of conversation, if the size of the memory block of the tensor of the operator meets the requirements of this round of conversation, the memory block can be directly used to store the data of the tensor, and there is no need to apply for the memory block of the tensor again. That is to say, in any round of conversation after the first round, a judgment is made for the tensor of each operator; if there is a memory block corresponding to a tensor, and the size of the memory block meets the storage requirements of the tensor in this round of conversation (the storage requirements are determined according to the dimension of the tensor), the existing memory block is reused to store the data of the tensor; if there is a memory block corresponding to a tensor, and the size of the memory block does not meet the storage requirements of the tensor in this round of conversation, the memory block corresponding to the tensor is released, and then a memory block of a new size is applied for according to the dimension of the tensor. By reusing the existing memory blocks, the number of memory applications and releases can be reduced, the processing speed of the model can be improved, the waiting time for the user to receive a reply after inputting a conversation can be reduced, and the user experience can be enhanced.

[0068] The method for dynamic memory management provided by the embodiment of the present application is applicable to application scenarios that require frequent memory application and release, and can reduce the number of memory applications and releases and improve the performance of the model.

[0069] It can be known that the embodiment of the present application is introduced by taking a large language model as an example, and it can also be applicable to other large models with a certain scale of the number of operators, such as image processing models, etc.

[0070] In some embodiments, in any round of conversation after the first round, based on the input tensor of the model in this round of conversation, the data transfer relationships between various operators in the model, the calculation logic and parameters of the operators, the dimensions of the tensors of each operator in the model are determined. Furthermore, based on the dimensions of the tensors of each operator, the required memory values of each operator are determined, that is, the storage requirements of the tensors of each operator in the model for this round of conversation are obtained. According to the magnitudes of the required memory values of the tensors of each operator in the model determined in this round of conversation, they are compared with the sizes of the memory blocks actually allocated for this operator in the conversations of previous rounds. If, in this round of conversation, the required memory value of the tensor of an operator is less than or equal to the memory block actually allocated for this operator in the conversation of the previous round, it is considered that the memory block actually allocated for the tensor of this operator in the conversation of the previous round meets its storage requirements for this round of conversation, and this memory block is reused. That is to say, the tensor of this operator does not need to allocate a new memory block in this round of conversation. If, in this round of conversation, the required memory value of the tensor of an operator is greater than the memory block actually allocated for the tensor of this operator in the conversation of the previous round, it is considered that the memory block actually allocated for the tensor of this operator in the conversation of the previous round cannot meet its storage requirements for this round of conversation. That is to say, it is necessary to release the memory block already allocated for the tensor of this operator in the conversation of the previous round in this round of conversation, and re-allocate a new memory block according to the required memory value of the tensor of this operator in this round of conversation.

[0071] It can be known that since the comparison is performed separately for the tensors of each operator in the model, after all operators have been compared, there may be one, some, or all tensors of operators that need to re-allocate memory blocks. Conversely, for the tensors of operators that do not need to re-allocate memory blocks, the memory blocks actually allocated for them in the conversations of previous rounds can be reused.

[0072] Exemplarily, the size of the memory actually allocated for the tensor of an operator needs to be greater than or equal to its required memory value. That is to say, the memory actually allocated for the tensor of an operator can be its required memory value, or the required memory value + a memory increment, where the memory increment is a positive increment.

[0073] In some embodiments, if the actually applied memory of the tensor of an operator is its required memory value + memory increment, then when determining whether the actually applied memory block of the tensor of the operator in the current round of conversation can meet its storage requirements in the current round of conversation, if the required memory value of the tensor of the operator determined in the current round of conversation is less than or equal to the required memory value + memory increment determined in the conversation of the previous round, it is considered that the actually applied memory block of the tensor of the operator in the conversation of the previous round meets its storage requirements in the current round of conversation. That is to say, the tensor of the operator does not need to re-apply for a memory block in the current round of conversation. If the required memory value of the tensor of the operator determined in the current round of conversation is greater than the required memory value + memory increment determined in the conversation of the previous round, it is considered that the actually applied memory block of the tensor of the operator in the conversation of the previous round cannot meet its storage requirements in the current round of conversation. That is to say, it is necessary to release the memory block already applied by the tensor of the operator in the conversation of the previous round in the current round of conversation and re-apply for a new memory block according to the required memory value + memory increment corresponding to the dimension of the tensor of the operator. It can be known that if an operator has multiple tensors, each tensor is judged separately.

[0074] Exemplarily, when applying for a memory block for the tensor of an operator that needs to re-apply for memory, the required memory value can be determined according to the dimension of the tensor of the operator, and a memory block of the corresponding size can be applied for according to the required memory value of the operator.

[0075] In another example, a memory block of the corresponding size can also be applied for according to the required memory value + memory increment of the tensor of the operator. This method is also called memory increment application.

[0076] Exemplarily, the memory increment application can be determined, for example, by the following formula (1) for the actually applied memory block of each tensor of an operator.

[0077]

[0078] Wherein, in the above formula (1), x refers to the required memory value corresponding to the dimension of the tensor of the operator, y refers to the actually applied memory value, and [log2x] represents taking the integer of the result of the logarithm of x to the base 2. As Figure 4 shown, if the required memory value corresponding to the dimension of the tensor of the operator is 10, then the actually applied memory value of the tensor of the operator after increment is 16. The 6 therein is the memory increment.

[0079] Exemplarily, the required memory value corresponding to the dimension of the tensor in the above formula (1) can be obtained by the following formula (2).

[0080] x = tensor_dimension × sizeof_data_type (Equation 2)

[0081] Among them, tensor_dimension represents the dimension of the tensor, and sizeof_data_type represents the memory space occupied by a single data in the tensor, that is, the data type in the above example. The unit of the obtained x value is Byte.

[0082] In another example, the memory increment application can also be determined by a fixed memory increment value. The fixed memory increment value is used to determine the size of the actual memory block applied for the tensor of each operator. For example, on the basis of the required memory value corresponding to the dimension of the tensor of the operator, a fixed memory increment value is added as the size of the actual memory value applied for the operator. The embodiments of the present application do not make too many restrictions.

[0083] In some embodiments, two rounds of a multi-round dialogue are taken as examples for illustration. Among them, the first round of dialogue is represented as the first-round dialogue, and the second round of dialogue is represented as any round of dialogue other than the first-round dialogue. As Figure 5 shown in (a) of, in the first round of dialogue, the input data of the model is: 10 tokens. That is to say, the number of word segments of the input data is 10, the required memory value of the dimension of the input tensor of the model is 10 × 4 = 40. Operator 1 is connected to the input of the model, and the required memory value corresponding to the dimension of tensor11 of Operator 1 is 40. Operator 2 is connected to Operator 1, and the dimension of tensor21 of Operator 2 is 25, and the corresponding required memory value is 100. After applying the memory increment according to the dimensions of tensor11 and tensor21 and the above formula (1) respectively, the actual memory value applied for tensor11 of Operator 1 is 64, and the increment is 24. The actual memory value applied for tensor21 of Operator 2 is 128, and the increment is 28.

[0084] As Figure 5As shown in (b) of , in the second round of conversation, the input data of the model is: 15 tokens. It can be known that the dimension of the input tensor of the model is 15. Operator 1 is connected to the input of the model. The dimension of tensor11 of Operator 1 is 15, and the corresponding required memory value is 60. Operator 2 is connected to Operator 1. The dimension of tensor21 of Operator 2 is 30, and the corresponding required memory value is 120. In the first round of conversation, the actual memory value allocated for tensor11 of Operator 1 is 64, which is greater than the required memory value of tensor11 of Operator 1 in this round of conversation. That is to say, the memory block allocated for tensor11 in the first round of conversation can meet its storage requirements in this round of conversation, and the memory block of tensor11 is reused in this round of conversation. In the first round of conversation, the actual memory value allocated for tensor21 of Operator 2 is 128, which is greater than the required memory value of tensor21 of Operator 2 in this round of conversation. That is to say, the memory block allocated for tensor21 in the first round of conversation can meet its storage requirements in this round of conversation, and the memory block of tensor21 is reused in this round of conversation.

[0085] In some embodiments, as Figure 6 shown in (a) of , in the first round of conversation, the input data of the model is: 10 tokens. The determination of the input tensor of the model and the tensors of each operator is similar to the above example. The required memory value of tensor11 of Operator 1 is 40, the dimension after increment is 64, and the increment is 24. The required memory value of tensor21 of Operator 2 is 100, the actual memory value allocated is 128, and the increment is 28. Details are not elaborated here.

[0086] such as Figure 6As shown in (b), in the second round of conversation, the input data of the model is 20 tokens, the dimension of the input tensor of the model is 20, the dimension of tensor11 of operator 1 is 20, and the corresponding required memory value is 80. The dimension of tensor21 of operator 2 is 60, and the corresponding required memory value is 240. In the first round of conversation, the actual memory application value of tensor11 of operator 1 is less than the required memory value of tensor11 of operator 1 in this round of conversation. That is to say, the memory block allocated for tensor11 in the first round of conversation cannot meet its storage requirements in this round of conversation. Release this memory block, and according to the dimension of tensor11 and the memory increment application method of the above formula (1), determine that the actual memory application value of tensor11 is 128, and re-allocate a memory block of the corresponding size according to the actual memory application value of tensor11. In the first round of conversation, the actual memory application value of tensor21 of operator 2 is 128, which is less than the required memory value of tensor21 of operator 2 in this round of conversation. That is to say, the memory block allocated for tensor21 of operator 2 in the first round of conversation cannot meet its storage requirements in this round of conversation. Release this memory block, and according to the dimension of tensor21 in this round of conversation and the memory increment application method of the above formula (1), determine that the actual memory application value of tensor21 is 256, and re-allocate a memory block of the corresponding size according to the actual memory application value of tensor21.

[0087] In some embodiments, as Figure 7 shown in (a), in the first round of conversation, the input data of the model is: 12 tokens. The determination of the input tensor of the model and the tensors of each operator is similar to the above example. The required memory value of tensor11 of operator 1 is 48, the actual memory application value is 64, and the increment is 16. The required memory value of tensor21 of operator 2 is 120, the actual memory application value is 128, and the increment is 8.

[0088] such as Figure 7As shown in (b) of , in the second-round conversation, the input data of the model is 15 tokens, the dimension of the input tensor of the model is 15, the dimension of tensor11 of operator 1 is 15, and the corresponding required memory value is 60. The dimension of tensor21 of operator 2 is 34, and the corresponding required memory value is 136. In the first-round conversation, the actually allocated memory value of tensor11 is greater than the required memory value of tensor11 in this round of conversation. That is to say, the memory block allocated for tensor11 in the first-round conversation can meet its storage requirements in this round of conversation, and this memory block is reused. In the first-round conversation, the actually allocated memory value of tensor21 is less than the required memory value of tensor21 in this round of conversation. That is to say, the memory block allocated for tensor21 in the first-round conversation cannot meet its storage requirements in this round of conversation. Release this memory block, and according to the dimension of tensor21 in this round of conversation and the memory increment application method of the above formula (1), allocate a memory block of the corresponding size for tensor21.

[0089] The above Figure 5 introduced an example where tensors of all operators in the model do not need to allocate new memory blocks. In this example, since tensors of all operators in the model do not need to allocate and release memory in this round of conversation, a large amount of processing time can be saved. Figure 6 introduced an example where tensors of all operators in the model need to allocate new memory blocks. In this example, since tensors of all operators in the model need to allocate and release memory in this round of conversation, this round of conversation requires more processing time. However, this is a necessary memory allocation and release process to avoid model crashes. Figure 7 introduced an example where tensors of some operators in the model need to allocate new memory blocks while tensors of other operators do not. In this example, for tensors of operators that need to allocate new memory blocks, new memory blocks are allocated. Compared with the existing memory allocation and release mechanism, the processing time of tensors of those operators that do not need to allocate new memory blocks is saved. Generally speaking, in multi-round conversations, the embodiments of the present application can effectively reduce the number of memory allocations and releases.

[0090] It can be known that in the embodiments of the present application, the dimensions of the tensors of each operator are judged separately. When one or some of these tensors need to reapply for memory, separate memory applications are made for these tensors. The numerical values in the above examples are only for illustrative purposes and are not specific limitations. In actual applications, the process of data operation and transfer between operators is relatively complex. The tensors of operators are determined by the input data, the data transfer relationship between each operator in the actual model, the parameters of the operator, and the calculation logic. The above examples only illustrate with one operator corresponding to one tensor. In actual applications, one operator can correspond to multiple tensors. The embodiments of the present application will not be elaborated here.

[0091] In some embodiments, the data transfer relationship between operators in the model, the parameters and calculation logic within the operator are relatively fixed. Therefore, the number of word segments of the input data has a greater influence on the dimensions of the tensors of the operator, and the dimensions of the tensors of the operator can be determined according to the dimensions of the input tensors of the model, the data transfer relationship between operators, the parameters and calculation logic within the operator. Exemplarily, the memory required for the dimensions of the input tensors of the model determined in the current round of conversation can be compared with the actual memory applied for this input tensor in the historical round of conversation. If the memory required for the input tensors of the model in the current round of conversation is less than the actual memory applied for this input tensor in the historical round of conversation, then it is no longer necessary to judge whether the tensors of the operators in the model need to reapply for memory blocks, and directly reuse the memory blocks applied for the tensors of the operators in the historical round of conversation.

[0092] In one example, as Figure 8 shown, the embodiments of the present application provide a large language model system 100. The large language model system 100 is deployed in a terminal device and includes a memory planning module 110, a judgment module 120, and a data reasoning module 130. Among them, the judgment module 120 is used to determine the tensors that need to reapply for memory after the memory planning module 110 determines the tensors of each operator in the model. The memory planning module 110 generates a memory application request according to the dimensions of the tensors that need to reapply for memory determined by the judgment module 120 and sends it to the memory allocator of the device. The memory allocator releases the memory blocks of the tensors that do not meet the need to reapply for memory according to the memory application request, and reapplies new memory blocks for these tensors that need to reapply for memory. After completing the memory application, the processor stores the data into the corresponding memory block according to the storage address, and the data reasoning module 130 reads the data from the memory block for a series of operations and transfers to obtain the output data.

[0093] Figure 9 A memory dynamic management method according to an embodiment of the present application is applied to a terminal device. The terminal device is deployed with a large language model system. In each round of multi-round conversations, the terminal device executes the following S901 to S909:

[0094] S901. Obtain the input data of the large language model system.

[0095] Exemplarily, the input data includes: word vectors obtained by performing word segmentation processing and encoding processing on the language text in the text conversation input by the user in the current round of conversation, and information representing the number of word segments. The information representing the number of word segments can be obtained by counting the number of word segments after performing word segmentation processing on the language text.

[0096] S902. Determine the input tensor of the model according to the input data.

[0097] Exemplarily, according to the input data, determine the information representing the number of word segments. When obtaining the number of word segments of the language text in the current round of conversation, the input tensor of the model can be determined according to the number of word segments.

[0098] S903. Determine whether the current round of conversation is the first round of conversation.

[0099] In some embodiments, an instruction can be called to check whether there is a memory block in the memory that has been applied for the tensor of the operator. If the call result indicates that there is no memory block in the memory that has been applied for the tensor of the operator, it is determined that this round of conversation is the first round of conversation. If the call result indicates that there is a memory block in the memory that has been applied for the tensor of the operator, it is determined that this round of conversation is not the first round of conversation.

[0100] S904. When it is determined that the current round of conversation is the first round of conversation, determine the dimension of the tensor of each operator in the model according to the dimension of the input tensor of the model, and generate a memory application request.

[0101] Exemplarily, the memory application request may include the size of the actually applied memory value of each operator in the model.

[0102] S905. When it is determined that the current round of conversation is not the first round of conversation, determine the tensors of the operators that need to re-apply for memory blocks in this round of conversation according to the dimensions of the tensors of each operator in the model and the memory blocks of the tensors of this operator applied in the previous historical rounds of conversation.

[0103] S906. For the tensors of the operators that do not need to re-apply for memory blocks, reuse the memory blocks applied for the tensors of this operator in the historical rounds of conversation.

[0104] S907. For the dimensions of the tensors of the operators that need to re-apply for memory blocks, release the memory blocks applied for the tensors of the operators in the historical round of conversation, and generate a memory application request.

[0105] Exemplarily, the memory application request may include the size of the actually applied memory value of the operator that needs to re-apply for a memory block.

[0106] S908. Allocate memory blocks according to the memory allocation request.

[0107] In some embodiments, after obtaining the memory application request, the memory allocator allocates corresponding-sized memory blocks for the operators of the model according to the dimensions of the actually applied tensors of the operators of the model included in the memory application request.

[0108] S909. Process the data in the memory block to generate output data.

[0109] Exemplarily, the processor of the terminal device reads the corresponding data from the memory block, and performs operations and transmissions according to the operation rules of the operator to generate output data.

[0110] In some embodiments, after a round of conversation ends, if no next input from the user is received within a preset waiting duration, in order to improve the utilization rate of memory, the memory block will be released. After the memory block is released, if the next input from the user is received, then this round of input is regarded as the first round of conversation, and the tensor dimensions of each operator of the model are determined according to the input data of the user, and memory blocks are applied for the tensors of each operator of the model. It can be known that regarding it as the first round of conversation does not mean that this round of conversation is the first round of conversation, but only means that it is necessary to re-apply for memory blocks according to the tensors of each operator in the model.

[0111] Exemplarily, after a round of conversation ends, the terminal device detects an instruction to exit the model, and then releases the memory block. The instruction to exit the model may be that the user closes the human-machine conversation, or the program has an interruption.

[0112] In some embodiments, as the number of dialogue turns increases and the number of word segments in the text input by the user increases, the size of the memory blocks of the tensors of each operator in the model tends to increase. In a certain round of dialogue, if the number of word segments in the text input by the user decreases significantly compared to the previous round of dialogue, and if the memory blocks allocated for the tensors of each operator in the previous round of dialogue are still reused, then there may be a lot of idle memory. To address this issue, exemplary, a memory reset threshold can be preset. In the current round of dialogue, if the dimension of the input tensor of the model is less than the preset memory reset threshold for the memory blocks allocated for the input tensor in the previous round of dialogue, release the memory blocks allocated for the tensors of each operator in the model and re-allocate new memory blocks according to the actual tensors required by the operators determined in this round of dialogue. This avoids the continuous increase in the size of the memory blocks occupied by the model, which may lead to a decrease in memory utilization. For example, in the first round of a multi-round dialogue, the number of word segments in the text input by the user is 10. The dimension of the input tensor of the model is 10, and the corresponding required memory value is 40. In the first round of dialogue of the large language model system, the actual memory value allocated for the input tensor of the model is 64. In the second round of dialogue, the dimension of the input tensor of the model is 20, and the corresponding required memory is 80. In the second round of dialogue of the large language model system, the actual memory value allocated for the input tensor of the model is 128. In the third round of dialogue, the dimension of the input tensor of the model is 4, and the corresponding required memory value is 16. At this time, if the large language model system reuses the memory block with a size of 128 determined in the second round of dialogue, it will cause waste of system memory resources.

[0113] It can be known that if the preset memory reset threshold is set too small, memory application and release will be triggered frequently, resulting in a slower processing speed of the model. If the preset memory reset threshold is set too large, it may cause serious waste of system memory resources. Therefore, the value of the preset memory reset threshold needs to balance the processing speed of the model while taking into account the utilization rate of memory resources. The specific value is set according to the actual situation and is not limited in the embodiments of this application.

[0114] An embodiment of the present application provides a method for dynamic memory management, which is applied to a terminal device. A large language model system is deployed on the terminal device. In each round of multi-round conversation, the following process is executed: obtain input data, determine the input tensor of the model according to the number of word segments in the input data. When it is determined that the current conversation is the first round of conversation, determine the dimension of the tensor of each operator in the model according to the input tensor of the model and the data transfer relationship between the operators in the model, the parameters and calculation logic within the operator, and apply for a memory block according to the dimension of the tensor of each operator. If it is determined that the current conversation is not the first round of conversation, obtain the memory blocks that have already been applied for the tensors of each operator in the model, determine the tensors that need to re-apply for memory according to the dimension of the tensors of the operators determined in this round of conversation and the memory blocks applied for the tensors of the operators in the historical round of conversation, and apply for new memory blocks according to the dimensions of the tensors that need to re-apply for memory. For the tensors that do not need to re-apply for memory, reuse the memory blocks applied for the tensors in the historical round of conversation. Thereby reducing the number of times of triggering memory application and release, improving the processing speed of the model, reducing the waiting time for the user to receive an answer, and enhancing the user experience.

[0115] The terminal device mentioned in the method for dynamic memory management provided by the embodiment of the present application may be a portable computer (such as a mobile phone), a tablet computer, a laptop computer, a personal computer (PC), a wearable terminal device (such as a smart watch), an augmented reality (AR) virtual reality (VR) device, a vehicle-mounted computer, a smart TV, etc., which can deploy a large language model. The following embodiments do not make special restrictions on the specific form of the terminal device.

[0116] Exemplarily, Figure 10 shows a schematic structural diagram of a terminal device 200. As Figure 10 shown, it shows a schematic structural diagram of a terminal device 200. The terminal device 200 may include a processor 210, an external memory interface 220, an internal memory 221, an audio module 230, a display screen 240, a communication module 250, a power module 260, an input device 270, a sensor module 280, etc. Among them, the sensor module 280 may include a touch sensor, etc.

[0117] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the terminal device 200. In other embodiments of the present application, the terminal device 200 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0118] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0119] Among them, the controller may be the nerve center and command center of the terminal device 200. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0120] The operating system of the terminal device 200 may run on the application processor, which is used to manage the hardware and software resources of the terminal device 200. For example, manage and configure the memory, determine the priority order of system resource supply and demand, control input and output devices, operate the network, manage the file system, manage the driver, etc. The operating system may also be used to provide an operation interface for users to interact with the system. Among them, various software, such as drivers, application programs, etc., may be installed in the operating system.

[0121] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may store the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0122] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0123] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the terminal device 200. In other embodiments of the present application, the terminal device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods in the above embodiments.

[0124] The external memory interface 220 may be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the terminal device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0125] The internal memory 221 can be used to store one or more computer programs, and the one or more computer programs include instructions. The processor 210 can execute the application running method provided in some embodiments of the present application, as well as various applications and data management, etc., by running the above instructions stored in the internal memory 221. The internal memory 221 can include a code storage area and a data storage area. Among them, the data storage area can store data created during the use of the terminal device 200. In addition, the internal memory 221 can include high-speed random access memory, and can also include non-volatile memory, such as one or more disk storage components, flash memory components, universal flash storage (UFS), etc. In some embodiments, the processor 210 can execute the application running method provided in the embodiments of the present application, as well as other applications and data management, by running the instructions stored in the internal memory 221 and / or the instructions stored in the memory provided in the processor 210.

[0126] The terminal device 200 can implement audio functions through the audio module 230, speakers, microphones, and application processors, etc. For example, music playback, recording, etc. The audio module 230 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 230 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 230 can be provided in the processor 210, or some functional modules of the audio module 230 can be provided in the processor 210.

[0127] In the embodiments of the present application, the audio module 230 can convert the user's voice into natural language text, or convert the text output by the model into voice.

[0128] The speaker, also called the "loudspeaker", is used to convert an audio electrical signal into a sound signal.

[0129] The microphone, also called the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. The user can input a sound signal into the microphone by speaking close to the microphone.

[0130] The communication function of the terminal device 200 can be implemented through the antenna 1, antenna 2, and communication module 250, etc.

[0131] The communication module 250 can provide solutions for wireless communications applied to the terminal device 200, including cellular, Wi-Fi, Bluetooth (BT), wireless data transmission modules (such as 433 MHz, 868 MHz, 915 MHz), etc. The communication module 250 can be one or more devices integrating at least one communication processing module. The communication module 250 receives electromagnetic waves via Antenna 1 or Antenna 2, filters and frequency-modulates the electromagnetic wave signals, and sends the processed signals to the processor 210. The communication module 250 can also receive the signals to be transmitted from the processor 210, frequency-modulate and amplify them, and convert them into electromagnetic waves through Antenna 1 or Antenna 2 for radiation.

[0132] The terminal device 200 realizes the display function through the GPU, the display screen 240, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 240 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information. In the embodiments of the present application, the GPU is mainly responsible for processing the operations of a large number of operators of the large language model.

[0133] The display screen 240 is used to display images, videos, etc. The display screen 240 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 200 may include 1 or N display screens 240, where N is a positive integer greater than 1. In the embodiments of the present application, the display screen 240 can be used to display the UI and receive user operations on the UI.

[0134] In some embodiments, a pressure sensor, a touch sensor, etc. are provided on the display screen 240. The pressure sensor is used to sense pressure signals and can convert the pressure signals into electrical signals. When a touch operation acts on the display screen 240, the terminal device 200 detects the intensity of the touch operation according to the pressure sensor. The terminal device 200 can also calculate the position of the touch according to the detection signal of the pressure sensor. The touch sensor, also called a "touch panel", can form a touch screen, also called a "touch screen", with the display screen 240. The touch sensor is used to detect touch operations acting on or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can also be provided through the display screen 240.

[0135] The power module 260 can be used to supply power to each component included in the terminal device 200. In some embodiments, the power module 260 can be a battery, such as a rechargeable battery.

[0136] The input device 270 can include a keyboard, a mouse, etc. The keyboard is used to input English letters, numbers, punctuation marks, etc. into the terminal device 200, thereby sending commands to the terminal device 200 and inputting data, etc. The mouse is an indicator for positioning the horizontal and vertical coordinates of the display system of the terminal device 200 and is used to input instructions to the terminal device 200. Among them, the input device 270 can be connected to the terminal device 200 by a wired connection method. For example, the input device 270 is connected to the terminal device 200 through a GPIO interface, a USB interface, etc. The input device 270 can also be connected to the terminal device 200 wirelessly. For example, the input device 270 is connected to the terminal device 200 through Bluetooth, infrared, etc.

[0137] The sensor module 280 of the terminal device 200 also includes a touch sensor. The touch sensor is used to detect the touch operation of the user on the screen of the terminal device. The terminal device generates natural language text according to the user's touch operation, and is used to perform a series of operations such as word segmentation, encoding, and reasoning on the natural language text in subsequent processing.

[0138] In the embodiments of the present application, the operating system of the above terminal device 200 can be the same as the system in the above example, or other systems with the same or similar display process as the process in the above example system, system, etc.

[0139] An embodiment of the present application provides a terminal device, which may include: a memory and one or more processors. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the terminal device can execute each function or step executed by the mobile phone in the above method embodiment. The structure of the terminal device may refer to Figure 10 the structure of the terminal device shown.

[0140] An embodiment of the present application also provides a chip system (for example, a system on a chip (SoC)), as Figure 11 shown, the chip system includes at least one processor 1101 and at least one interface circuit 1102. The processor 1101 and the interface circuit 1102 can be interconnected by a line. For example, the interface circuit 1102 can be used to receive signals from other devices (such as the memory of the terminal device). For another example, the interface circuit 1102 can be used to send signals to other devices (such as the processor 1101 or the camera of the terminal device). Exemplarily, the interface circuit 1102 can read the instructions stored in the memory and send the instructions to the processor 1101. When the instructions are executed by the processor 1101, the terminal device can execute each step in the above embodiment. Of course, the chip system can also include other discrete devices, and the embodiments of the present application do not make specific limitations thereto.

[0141] An embodiment of the present application also provides a computer-readable storage medium, including computer instructions. When the computer instructions run on the terminal device, the terminal device executes each function or step executed by the terminal device 200 in the above method embodiment.

[0142] An embodiment of the present application also provides a computer program product. When the computer program product runs on the terminal device, the computer executes each function or step executed by the terminal device 200 in the above method embodiment. For example, the computer can be the above terminal device 200.

[0143] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0144] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0145] The units described as separate components may or may not be physically separated. The components displayed as units can be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0147] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.

[0148] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for dynamic memory management, characterized in that, Applied to a terminal device, the terminal device is deployed with a target model, the target model includes a first operator, and the first operator is any operator in the target model. The method includes: Obtain first input data of the target model; If it is determined that there is a first memory block and the size of the first memory block meets the requirement for storing first data, use the first memory block to store the first data; the first data is the data generated by the first operator processing the first input data, the first memory block is the memory allocated for the first operator during the target model's processing of second input data, and the size of the first memory block is determined according to the number of word segments of the second input data.

2. The method according to claim 1, wherein The method further includes: If it is determined that there is a first memory block and the size of the first memory block does not meet the requirement for storing the first data, release the first memory block and apply for a second memory block according to the first memory value required for storing the first data; Use the second memory block to store the first data.

3. The method according to claim 1 or 2, characterized in that, Before obtaining the first input data of the target model, the method further includes: Obtain the second input data of the target model; Determine the dimensions of the input tensor and output tensor of the first operator of the target model according to the number of word segments of the second input data; Apply for the first memory block according to the dimensions of the input and output tensors of the first operator of the target model.

4. The method according to claim 1, characterized in that, The determination that the size of the first memory block meets the requirement for storing first data includes: If the number of word segments of the first input data is less than the number of word segments of the second input data, determine that the size of the first memory block meets the requirement for storing first data.

5. The method according to claim 2, characterized in that The determination that the size of the first memory block does not meet the requirement for storing first data includes: If the first memory value is greater than the size of the first memory block, determine that the size of the first memory block does not meet the requirement for storing first data.

6. The method according to claim 1, wherein After obtaining the first input data of the target model, the method further includes: If it is determined that there is no memory allocated for the first operator, determine the dimensions of the input tensor and output tensor of the first operator of the target model according to the number of word segments of the first input data; Apply for the second memory block according to the dimensions of the input tensor and output tensor of the first operator; Use the second memory block to store the first data.

7. The method according to claim 2 or 6, characterized in that The method further includes: Determine a first memory value according to the dimensions of the input tensor and output tensor of the first operator and the correspondence between the dimensions of the tensor and the memory size; Determine the size of the second memory block according to the first memory value.

8. The method according to claim 7, wherein The determination of the size of the second memory block according to the first memory value includes: The size of the second memory block is the first memory value plus a fixed memory increment value.

9. The method according to claim 7, wherein The determination of the size of the second memory block according to the first memory value includes: Where x is the first memory value and y is the second memory value; Determine the size of the second memory block according to the second memory value.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: Determine the input tensor of the target model according to the number of word segments of the first input data; Determine the dimensions of the input tensor and output tensor of the first operator according to the dimensions of the input tensor of the target model, the data transfer relationship between the first operators of the target model, the parameters and calculation logic within the first operator.

11. The method according to any one of claims 1-10, characterized in that, The method further includes: Determine the first output of the target model according to the first input data.

12. The method according to claim 11, wherein After determining the first output of the target model according to the first input data, the method further includes: If no input data of the target model is received within a preset time period, release the memory blocks allocated for the input tensor and output tensor of each first operator in the target model.

13. A terminal device, characterized in that, The terminal device includes: a processor and a memory, the processor is coupled with the memory; the memory is used for storing computer program code; the computer program code includes computer instructions, when the processor executes the above computer instructions, the terminal device is caused to execute the method according to any one of claims 1-12.

14. A computer-readable storage medium, characterized in that, Including computer instructions, when the computer instructions run on a terminal device, the terminal device is caused to execute the method according to any one of claims 1-12.

15. A computer program product, characterized in that, When the computer program product runs on a terminal device, the terminal device is caused to execute the method according to any one of the above claims 1-12.

Citation Information

Patent Citations

  • Memory management method and device in neural network forward calculation process

    CN108829610A

  • Memory management method and device, mobile terminal and storage medium

    CN109815162A

  • Deployment method of voice processing model, electronic equipment and storage medium

    CN115080240A

  • Deep learning memory allocation optimization method and system

    CN116302461A

  • Memory management method and device for neural network model, equipment, medium and product

    CN116893904A