Task processing method and apparatus therefor
By splitting the prompt information into two parts: business scenario and intention, and directly obtaining feature information from the cache area for inference processing, the problem of large amount of computing in the long prompt information sequence of AI models is solved, and processing efficiency is improved.
Patent Information
- Application Number
- PCT/CN2025/078792
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2025-02-24
- Publication Date
- 2025-09-04
AI Technical Summary
When existing AI models process long prompt information sequences, the calculation amount is large, resulting in low processing efficiency and long time.
The prompt information is split into two parts indicating the business scenario and intention, and the characteristic information of the business scenario is directly obtained from the cache area, and the characteristic information of the intention is calculated through the AI model for inference processing.
It reduces the amount of computing in the task processing process of AI models, improves processing efficiency, and reduces waiting time.
Smart Images

Figure CN2025078792_04092025_PF_FP_ABST
Abstract
Description
Task processing method and device
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202410230371.2 filed on February 29, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present invention relates to the field of artificial intelligence, and specifically to a task processing method and device thereof. Background Art
[0004] With the development of artificial intelligence (AI) technology, AI models for processing various tasks have been widely used.
[0005] Currently, when AI models are used to process tasks such as document summarization and knowledge question-answering, in the case of a sequence of prompt information input by the user, it is necessary to perform matrix operations on the information sequence in the pre-filling stage to obtain the feature information of the information sequence, and then perform inference processing based on the feature information to obtain the inference result. However, when the sequence of prompt information input by the user is long, the amount of computation required to perform matrix operations on the information sequence is large, which makes the entire calculation process take a long time, often taking hundreds of milliseconds or even seconds to generate an inference result. As a result, the processing process of the AI model takes a long time, resulting in low processing efficiency. Summary of the Invention
[0006] The purpose of the embodiments of the present application is to provide a task processing method and device thereof, which can reduce the amount of computation of the AI model during task processing and improve processing efficiency.
[0007] In the first aspect, an embodiment of the present application provides a task processing method, which includes: obtaining first prompt information of a first task, and splitting the first prompt information into first sub-information and second sub-information, the first sub-information indicating the business scenario of the first task, and the second sub-information indicating the intention of the first task; obtaining first feature information of the first sub-information from a first cache area, the first cache area including at least one feature information, each feature information corresponding to a type of business scenario; performing inference processing on the first feature information and the second feature information through a first AI model to obtain a first inference result of the first task, and the second feature information is obtained by transforming the second sub-information.
[0008] In the second aspect, an embodiment of the present application provides a task processing device, which includes: an acquisition module and a processing module, wherein: the acquisition module is used to obtain first prompt information of a first task, and split the first prompt information into first sub-information and second sub-information, the first sub-information indicates the business scenario of the first task, and the second sub-information indicates the intention of the first task; the acquisition module is also used to obtain first feature information of the first sub-information from a first cache area, the first cache area includes at least one feature information, and each feature information corresponds to a type of business scenario; the processing module is used to perform inference processing on the first feature information and the second feature information through a first AI model to obtain a first inference result of the first task, and the second feature information is obtained by transforming the second sub-information.
[0009] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0010] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0011] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.
[0012] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the method described in the first aspect.
[0013] In an embodiment of the present application, first prompt information for a first task is obtained and split into first sub-information and second sub-information, wherein the first sub-information indicates a business scenario for the first task, the second sub-information indicates a business scenario for the first task, and the second sub-information is used to indicate the intent of the first task. First feature information of the first sub-information is obtained from a first buffer, wherein the first buffer includes at least one feature information, each feature information corresponding to a type of business scenario. The first feature information and the second feature information are inferred by a first AI model to obtain a first inference result for the first task, and the second feature information is obtained by transforming the second sub-information. Through this method, when performing task processing, by splitting the prompt information into information indicating the business scenario and information indicating the schematic diagram, then directly obtaining feature information of the information indicating the business scenario from the buffer, calculating feature information of the information indicating the schematic diagram in the prompt information, and then performing inference processing based on the feature information of these two parts of information, there is no need to transform the entire prompt information to obtain the feature information of the entire prompt information. Instead, only a portion of the information indicating the schematic diagram in the prompt information needs to be transformed to obtain the feature information of the information, thereby reducing the amount of computation of the AI model during task processing and improving processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] FIG1 is a schematic diagram of a large model processing process in the related art;
[0015] FIG2 is a flowchart of a task processing method provided in an embodiment of the present application;
[0016] FIG3 is a schematic diagram of cache information of a cache area provided in an embodiment of the present application;
[0017] FIG4 is a second schematic diagram of cache information of a cache area provided in an embodiment of the present application;
[0018] FIG5 is a schematic diagram of a task processing device provided in an embodiment of the present application;
[0019] FIG6 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0020] FIG7 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0022] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0023] The terms "at least one" and "at least one of" in the specification and claims of this application refer to any one, any two, or a combination of more than two of the objects included. For example, at least one of a, b, and c can be represented by: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two" means two or more, and its meaning is similar to "at least one".
[0024] The following is an explanation of the nouns and terms involved in the embodiments of the present application.
[0025] Large Language Model (LLM): Large language models generally refer to deep learning models with large parameter sizes and the ability to learn complex language representations. These models are typically initialized through pre-training and then fine-tuned to adapt to specific tasks.
[0026] Traditional natural language processing systems can be limited by limited rules and corpora when faced with complex language understanding and generation tasks. To overcome these limitations, large language models have been widely used in recent years. Large language models are deep learning-based models with a large number of parameters, enabling them to learn and understand large amounts of natural language data.
[0027] Prompt: A prompt is an instruction or question provided by the user or system to the artificial intelligence model to guide the model to generate corresponding output or perform a specific task.
[0028] Token: In large language models, a "token" generally refers to the smallest unit of text processed by the model. This can be a character, a word, a subword, or other larger text units. These tokens are the basic units processed by the model during training and inference.
[0029] When processing tasks using a large language model, the input prompt word (Pompt) is encoded into a token. The large model's network structure and parameters continuously generate the next token with the highest probability based on all previous tokens, and finally decode it into text. The inference process of a generative model is as follows: given a text input, the model outputs an answer (of length N), performing N inference cycles. This means that this type of model outputs only one token per inference. This output token is concatenated with the input tokens and used as input for the next inference cycle, repeating this process until a terminator is encountered. This creates a problem: the input tokens for each inference cycle become longer, resulting in increased inference computing power. In most business scenarios (such as document summarization, knowledge question answering, and intent classification), the prompt word length for a complete conversation can exceed 1,000 tokens, placing high demands on computing power and performance.
[0030] The large model processing process includes a pre-filling phase and a decoding phase. The pre-filling phase occurs during the calculation of the first output token. The calculation requires calculating and saving the key and value of the input token. The calculated key and value are then input into the Transformer layer of the large model for processing. When the input token is long, the long information needs to be converted into a vector for matrix operations. During the operation, there will be a large number of matrix multiplication (gemm) operations, which is slow to process, resulting in a long processing time for this stage. The decoding phase occurs during the calculation of the second output token to the last token. Each inference uses the input token and output token of the previous inference as the input of the current inference to obtain the current output token.
[0031] Figure 1 is a schematic diagram of the large model processing process in the related art. As shown in Figure 1, in the pre-filling stage, the text "2048, 918, 1216, 1925" is input into the large model for reasoning, and the first reasoning outputs the reasoning result (i.e., the reasoning result) "9220". Then, in the decoding stage, the output result "9220" of the first reasoning and the input "2048, 918, 1216, 1925" of the first reasoning are used as the input of the second reasoning for reasoning, and the reasoning result "817" is output. Then, the output of the second reasoning and the input of the second reasoning are used as the input of the third reasoning for reasoning, and the reasoning result "0809" is output. Similarly, the output of the third reasoning and the input of the third reasoning are used as the input of the fourth reasoning for reasoning, and the reasoning result "999" is output, thereby obtaining the final reasoning result "9220, 817, 0809, 999".
[0032] The task processing method provided in the embodiment of the present application can be applied to the scenario of knowledge question answering using a large language model.
[0033] The task processing method provided in the embodiment of the present application, when processing a knowledge question and answer task, assumes that the user inputs a prompt message: "You are a professional translation assistant, please translate the following sentence into English: How is the weather today?", the task processing device splits the prompt message into information indicating the business scenario "You are a professional translation assistant, please translate the following sentence into English" and information indicating the schematic diagram "How is the weather today?", and then directly obtains the feature information of "You are a professional translation assistant, please translate the following sentence into English" from the cache area, and calculates the feature information of "How is the weather today?", and then performs inference processing based on these two parts of feature information to obtain the inference result of the knowledge question and answer task, so that there is no need to transform the entire prompt message to obtain the feature information of the entire prompt message, but only needs to transform the part of the information indicating the schematic diagram in the prompt message to obtain the feature information of the part of the information, thereby reducing the amount of calculation in the processing process and improving processing efficiency.
[0034] The task processing method provided in the embodiments of the present application can be executed by an electronic device or at least one of a functional module and an entity module in the electronic device that can implement the task processing method. The specific execution subject can be determined based on actual usage requirements and is not limited by the embodiments of the present application. The following embodiments illustrate the task processing method provided in the embodiments of the present application by taking a task processing device executing the task processing method as an example.
[0035] The task processing method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0036] FIG2 is a flowchart of a task processing method provided in an embodiment of the present application. As shown in FIG2 , the task processing method may include the following steps S201 to S203:
[0037] Step S201: The task processing apparatus obtains first prompt information of a first task, and splits the first prompt information into first sub-information and second sub-information.
[0038] The first sub-information indicates the business scenario of the first task, and the second sub-information indicates the intention of the first task.
[0039] In some embodiments of the present application, the above-mentioned first task can be tasks such as text generation, document summarization, intelligent question and answer, machine translation, question and answer classification, etc., which are not limited in the embodiments of the present application.
[0040] It is understandable that the first task may be a task that needs to be processed by an AI model or algorithm, and the first task may also be understood as a business.
[0041] In some embodiments of the present application, the first prompt information may be input information for triggering or guiding the large model to generate content.
[0042] It is understandable that inputting prompt information is essentially to improve the quality of model generation results by optimizing the prompt information, and inputting different prompt information may generate different results.
[0043] In some embodiments of the present application, the prompt information may be text, symbols, pictures, audio or video, etc. For example, the first prompt information is text input by the user, that is, a prompt word.
[0044] In some embodiments of the present application, the business scenario of the first task may be a business scenario for executing the first task. For example, the business scenario may be a translation scenario, an information query scenario, an intelligent question-and-answer scenario, a text generation scenario, or the like.
[0045] In some embodiments of the present application, the intention of the first task may be a specific intention to perform the first task. For example, the intention may be to translate the word "apple" into English, or to query the weather on January 19, or to write an article.
[0046] It is understandable that for tasks of the same type, their business scenarios are usually fixed or similar, but the specific intentions are usually changing.
[0047] In some embodiments of the present application, the first sub-information may be information related to the business scenario. The second sub-information may be information related to the intent. For example, the first sub-information may be information related to the business scenario, such as the task type, persona description, execution instructions, or restrictions. For example, the restrictions may be output format restrictions or output type restrictions. For example, the second sub-information may be specific query content.
[0048] For example, assuming that the prompt message input by the user is "You are a professional translation assistant, please translate the following sentence into English: How is the weather today?", the information describing the business scenario in the prompt message is "You are a professional translation assistant, please translate the following sentence into English:" (i.e., the first sub-information), and the information describing the intention in the prompt message is "How is the weather today?" (i.e., the second sub-information).
[0049] For example, assuming that the prompt message entered by the user is "Please help me check the prices of the following products on the shopping platform: check the prices of apples and strawberries.", then the information describing the business scenario in the prompt message is "Please help me check the prices of the following products on the shopping platform:" (that is, the first sub-information), and the information describing the intention in the prompt message is "Check the prices of apples and strawberries?" (that is, the second sub-information).
[0050] In some possible implementations, after obtaining the first prompt information of the first task, the task processing device may split the first prompt information into first sub-information and second sub-information according to a preset template.
[0051] In some embodiments of the present application, the aforementioned preset templates may include a general description format for business scenarios and a general description format for intents. For example, after receiving the input prompt information, the task processing device matches the content of the prompt information with the preset templates, determines the sub-information in the prompt information that matches the business scenario and the sub-information that matches the intent, and then splits the prompt information into these two sub-information parts.
[0052] In some embodiments of the present application, the preset template may be a text template, an image template, an audio template, a video template, or the like.
[0053] In some embodiments of the present application, corresponding templates can be preset for different types of tasks. For example, a translation task can be preset with a translation template, and a question-and-answer task can be preset with a question-and-answer template. This allows the task processing device to accurately split the prompt information for each type of task based on the template of the task.
[0054] In some possible implementations, after the task processing device obtains the first prompt information of the first task, it can split the first prompt information into first sub-information and second sub-information according to preset symbols.
[0055] Exemplarily, the preset symbol may be “:”, “.”, or “-”. Exemplarily, after receiving the input prompt information, the task processing device searches for the “:” in the prompt information, divides the information before the “:” into sub-information describing the business scenario, and divides the information after the “:” into sub-information describing the intention, and then splits the prompt information into these two parts of sub-information.
[0056] In some possible implementations, after obtaining the first prompt information of the first task, the task processing device may perform content analysis on the prompt information, and split the first prompt information into first sub-information and second sub-information according to the content analysis result.
[0057] Exemplarily, after obtaining the input prompt text, the task processing device performs semantic analysis on the prompt text to obtain the business scenario information and intention information in the prompt information, as well as the text content corresponding to the business scenario information and the text content corresponding to the intention information, and then splits the prompt information into the text content corresponding to the business scenario (i.e., the first word information) and the text content corresponding to the intention (i.e., the second sub-information).
[0058] In some embodiments of the present application, the task processing device may define fields corresponding to two sub-information obtained by splitting the prompt information (ie, the first prompt information), and the fields are used for internal processing procedures.
[0059] For example, the field corresponding to the first sub-information can be represented as system_prompt, and the field corresponding to the second sub-information can be represented as user_prompt. For example, system_prompt is the part of the input prompt information that remains unchanged when the user queries (query), and user_prompt is the part that changes with the user query.
[0060] It can be understood that a field can represent a variable associated with an object or class, and each field can be used to query the associated object or class.
[0061] For example, the input prompt is "You are a professional translation assistant, please translate the following sentence into English: How is the weather today?". The split prompt includes the prompt corresponding to the system_prompt field and the prompt corresponding to the user_prompt field, which are respectively expressed as "system_prompt: You are a professional translation assistant, please translate the following sentence into English:" and "user_prompt: How is the weather today?".
[0062] Step S202: The task processing device obtains first characteristic information of the first sub-information from the first buffer area.
[0063] The first cache area includes at least one piece of feature information, and each piece of feature information corresponds to a type of business scenario.
[0064] In some embodiments of the present application, the first characteristic information of the first sub-information may be a key (Key) and a value (Value) calculated based on the first sub-information.
[0065] For ease of description, in the embodiment of the present application, the key of the first sub-information is represented by K, and the value of the first sub-information is represented by V.
[0066] It should be noted that K is a weight index. By multiplying the attention index K (key) of other words with the attention weight (Query) of A, we can obtain the weighted attention of B on A. V (value) can be understood as the word vector in the current training corpus. It is the word vector obtained after intensive training using the current training corpus based on the original word vector. Q (query) can be understood as the attention weight of word vector A in the current training corpus. It stores the relationship between the remaining 99 words and A.
[0067] It should be noted that the feature information in the embodiments of the present application can be intermediate information obtained by processing the input prompt information when performing inference processing through the AI model. The AI model can perform inference processing based on this intermediate information and output an inference result. Because the existing process takes a long time to calculate this intermediate information, the entire processing flow takes a long time. In the embodiments of the present application, the feature information of the general information describing the business scenario is cached in the first cache area without the need for recalculation, thereby improving the processing speed.
[0068] In some embodiments of the present application, at least one feature information is pre-stored in the first buffer area.
[0069] In some embodiments of the present application, the first cache area may be a predefined cache area. Exemplarily, the first cache area may be represented as sys_kv_map.
[0070] In some embodiments of the present application, the first cache area may further store at least one first information, where the at least one characteristic information is characteristic information of at least one first information, and one first information corresponds to one characteristic information.
[0071] In some embodiments of the present application, the first cache area may be a cache area dedicated to storing characteristic information of the first information; or, the first cache area may be a cache area dedicated to storing the first information and the characteristic information of the first information; or, the first cache area may be a cache area for storing the first information, characteristic information corresponding to the first information, and other information.
[0072] In some embodiments of the present application, when the first information is stored in the first cache area, a hash value (ie, ghash) of the first information may be stored in the first cache area.
[0073] In some embodiments of the present application, in the case of first information and characteristic information of the first information cached in the first cache area, the task processing device stores the first information and the characteristic information of the first information in an associated manner in a key-value pair format, and uses the first information as the key and the characteristic information of the first information as the value.
[0074] It should be noted that the key-value pair here has a different meaning from the key and value of the sub-information (such as the first sub-information) in the embodiment of the present application. The key here is used as an index for querying feature information.
[0075] In some embodiments of the present application, the first information may be information related to a business scenario. For example, the first information may be information used to describe or indicate a business scenario.
[0076] It can be understood that, since the first information is information related to the business scenario, the characteristic information of the first information can correspond to a type of business scenario.
[0077] The following describes the information cached in the first cache area in the embodiment of the present application with reference to the accompanying drawings.
[0078] As shown in Figure 3, the first cache area stores the hash value ghash1 of information 1, the hash value ghash2 of information 2, and the hash value ghash3 of information 3. The first cache area also includes feature information KV cache1 of information 1, feature information KV cache2 of information 2, and feature information KV cache3 of information 3.
[0079] It should be noted that the embodiment of the present application only uses the storage of 3 pieces of information and 3 pieces of characteristic information in the cache area for illustrative purposes. The amount of information that can be stored in the first cache area can be set according to actual needs, and the embodiment of the present application does not limit this.
[0080] In some embodiments of the present application, before inputting the first prompt information into the model calculation, the task processing device first determines whether the first sub-information in the first prompt information exists in the first cache area. If so, the pre-filling stage calculation is skipped and the characteristic information of the first sub-information cached in the first cache area is directly used.
[0081] Step S203: The task processing device performs inference processing on the first feature information and the second feature information through the first AI model to obtain a first inference result of the first task.
[0082] The second characteristic information is obtained by transforming the second sub-information.
[0083] In some embodiments of the present application, the first AI model may be an LLM model or other large language model.
[0084] In some embodiments of the present application, the task processing device may perform inference processing on the first feature information and the second feature information through the Transformer layer of the first AI model to obtain a first inference result of the first task.
[0085] In some embodiments of the present application, the second characteristic information is characteristic information of the second sub-information, and the characteristic information includes a key and a value.
[0086] In some embodiments of the present application, the task processing device may convert the second sub-information into a vector, multiply the converted vector by the weight matrix of the calculated key to obtain the key of the second sub-information, and multiply the converted vector by the weight matrix of the calculated value to obtain the value of the second sub-information.
[0087] The following is an exemplary description of the process of calculating the second characteristic information with reference to the formula.
[0088] For example, the weight matrix of the calculation key is W K , calculate the key weight matrix W V , perform KV calculation on the information with a text sequence length of n (i.e., the second sub-information), assuming that the i-th text is x i The process of calculating the key is shown in the following formula (1), and the process of calculating the value is shown in the following formula (2): i =x i W K (1) v i =x i W V (2)
[0089] Among them, the key and value are represented by k respectively. i , v i .
[0090] In some embodiments of the present application, the task processing device can merge the first feature information and the second feature information to obtain merged feature information, and then perform inference processing on the merged feature information through the Transformer layer to obtain a first inference result.
[0091] In some embodiments of the present application, after obtaining the first inference result, the task processing device stores the first inference result in a cache area.
[0092] In some embodiments of the present application, a predictive decoding stage is performed, that is, the process from the second inference result to the last inference result outputted. At this time, each round of prediction only needs to read the input and output of the previous prediction from the cache, and append the new Key and Value calculated currently to the cache, and output the generated text word by word until a terminator is generated to end the entire inference process.
[0093] The task processing method provided in the embodiment of the present application is that when the task processing device performs task processing, the task processing device splits the prompt information into two parts: information indicating the business scenario and information indicating the schematic diagram. The task processing device then directly obtains the characteristic information of the information indicating the business scenario from the cache area, calculates the characteristic information of the partial information indicating the schematic diagram in the prompt information, and then performs inference processing based on the characteristic information of the two parts of information. Therefore, there is no need to transform the entire prompt information to obtain the characteristic information of the entire prompt information. Instead, only the partial information indicating the schematic diagram in the prompt information needs to be transformed to obtain the characteristic information of the partial information, thereby reducing the amount of calculation in the processing process and improving processing efficiency.
[0094] In some embodiments of the present application, the above-mentioned step S202 may include the following steps S202a and S202b.
[0095] Step S202a: The task processing device matches the first sub-information with at least one first information in the first buffer area.
[0096] Step S202b: The task processing device determines the characteristic information in the at least one first information that matches the first sub-information as the first characteristic information.
[0097] In some embodiments, when at least one first information is stored in the first cache area, the task processing device matches the first sub-information with the at least one first information one by one, and uses the characteristic information of the first information matched with the first sub-information as the characteristic information of the first information.
[0098] It should be noted that the first information matching the first sub-information may include: first information that is identical to the first sub-information or first information whose degree of similarity to the first sub-information is greater than or equal to a similarity threshold.
[0099] For example, taking the first sub-information as "You are a professional translation assistant, please translate the following sentence into English" as an example, the task processing device obtains information with the same or high degree of similarity to the above information content from the first cache area, and then obtains the Key and Value of the information as the characteristic information of the first sub-information, that is, the Key and Value.
[0100] In an embodiment of the present application, the task processing device can obtain characteristic information of information matching the first sub-information from the first cache area, and use the characteristic information as the characteristic information of the first sub-information without calculating the characteristic information of the first sub-information, thereby being able to quickly obtain the characteristic information of the first sub-information.
[0101] In some embodiments of the present application, before the above step S201, the task processing method provided in the embodiment of the present application may include the following steps S204 to S206:
[0102] Step S204: the task processing apparatus obtains the second prompt information of the second task, and splits the second prompt information of the second task into third sub-information and fourth sub-information.
[0103] The third sub-information indicates the business scenario of the second task, and the fourth sub-information indicates the intention of the second task.
[0104] Step S205: The task processing device transforms the third sub-information according to the transformation matrix to obtain third feature information.
[0105] Step S206 : When the query frequency of the third sub-information is greater than or equal to the first threshold, the task processing device stores the third sub-information and the third characteristic information in the first buffer area.
[0106] In some embodiments, the second task may be any task that the task processing device needs to process.
[0107] It should be noted that the explanation of the above step S204 can be found in the description of the above step S201, which will not be repeated here.
[0108] In some embodiments, the task processing device may convert the third sub-information into a vector, multiply the converted vector by the weight matrix of the calculated key to obtain the key of the third sub-information, and multiply the converted vector by the weight matrix of the calculated value to obtain the value of the third sub-information.
[0109] In some embodiments, after calculating the characteristic information of the third sub-information, the task processing device can obtain the query frequency of the third sub-information. If the query frequency is greater than or equal to the first threshold, the third sub-information and the third characteristic information are stored in the first cache area. If the query frequency is less than the first threshold, the third sub-information and the third characteristic information are not stored.
[0110] In some embodiments, the task processing device may set access frequency monitoring for each sub-information indicating the business scenario of the task (ie, system_prompt), and store the access frequency of each sub-information in the second cache area (sys_prompt_rate).
[0111] Exemplarily, the task processing device monitors the access count statistics n and the sequence length s of each sub-information within a preset time period (eg, 1 minute), and saves the access count and sequence length to the global cache sys_prompt_rate.
[0112] In some embodiments, the task processing device may associate and store the sub-information and the number of times the sub-information is accessed (i.e., the query frequency) and the sequence length in the form of a key-value pair in a global cache, and update the stored number of times the sub-information is accessed after the number of times the sub-information is accessed changes, so that the latest number of times the sub-information is accessed can be obtained from the global cache.
[0113] Figure 4 is a schematic diagram of the hash structure of sub-information, access counts, and sequence lengths stored in the global cache. As shown in Figure 4, the global cache stores the hash value ghash1 of sub-information 1, whose access count and sequence length are n1 and s1, respectively. It also stores the hash value ghash2 of sub-information 2, whose access count and sequence length are n2 and s2, respectively. The cache also stores the hash value ghash3 of sub-information 3, whose access count and sequence length are n3 and s3, respectively.
[0114] In some embodiments, when the query frequency of the third sub-information is greater than or equal to a first threshold, the task processing device associates the third sub-information and the third feature information in a key-value pair format and stores them in the first cache area.
[0115] Exemplarily, the third sub-information in the first buffer area may be expressed as: sys_kv_map.key=ghash(system_prompt), where sys_kv_map.key represents the third sub-information cached in the first buffer area, and system_prompt represents the third sub-information.
[0116] For example, the sequence length of the third sub-information is n, and the third characteristic information in the first buffer area can be expressed as: sys_kv_map.value = {(k1,v1),(k2,v2),…,(k n ,v n )}, where k1 and v1 represent the key and value of the first token (such as a character) in the sequence, respectively.
[0117] In an embodiment of the present application, the task processing device can store the characteristic information of sub-information whose query frequency is greater than or equal to the first threshold in the cache area, so that the characteristic information of information with a higher request frequency can be stored according to the user request frequency, thereby enabling efficient use of GPU cache resources.
[0118] In some embodiments of the present application, before storing the third sub-information and the third feature information in the first buffer in step S206, the task processing method provided in the embodiment of the present application may include the following step S207:
[0119] Step S207: when the amount of information cached in the first cache area is greater than or equal to the second threshold, the second information and the characteristic information of the second information in the first cache area are deleted.
[0120] The second information is: sub-information with the lowest query frequency or a query frequency less than a third threshold.
[0121] In some embodiments, the cached information in the first cache area may be all the information stored in the first cache area; or, the cached information in the first cache area may be at least one first information (system_prompt) stored in the first cache area; or, the cached information in the first cache area may be at least one first information stored in the first cache area and characteristic information of the first information.
[0122] In some embodiments, the data volume of the cached information may be the sum of the lengths of the cached information, or the sum of the data sizes of the cached information. The data volume of the cached information may be represented as cache_seq_count.
[0123] In some embodiments, the second threshold may be a predefined upper limit of the cache length, and the second threshold may be expressed as MAX_SEQ_LIMIT.
[0124] In some embodiments, before storing the third sub-information and the third characteristic information in the first cache area, the task processing device may calculate the length of the cached system_prompt in the first cache area. If the length is greater than or equal to the upper limit of the cache length, the second information and the characteristic information of the second information with a lower query frequency in the first cache area are deleted, and then the third sub-information and the third characteristic information are stored in the first cache area; or, if the length is less than the upper limit of the cache length, the third sub-information and the third characteristic information are directly stored in the first cache area.
[0125] In an embodiment of the present application, the task processing device can obtain the data size of the cached information in the cache area before storing the characteristic information of the sub-information with a query frequency greater than or equal to the first threshold in the cache area, and delete the information with a lower access frequency and the characteristic information of the information in the cache area when the data size of the cached information is greater than or equal to the cache upper limit value, so as to adaptively pre-process the cache according to the user request frequency, thereby efficiently using the GPU cache resources.
[0126] In some embodiments of the present application, after the third sub-information and the third feature information are stored in the first buffer in S206, the task processing method provided in the embodiment of the present application may include the following step S208:
[0127] Step S208: Update the first threshold to the query frequency of the third information.
[0128] The third information is the information with the lowest query frequency in the first cache area.
[0129] In some embodiments, after storing the third sub-information and the third characteristic information in the first cache area, the task processing device may update the first threshold to the query frequency of the information with the lowest query frequency in the first cache area.
[0130] Exemplarily, the first threshold is 10, and the query frequency of the information with the lowest query frequency in the first cache area is 12 times, then the first threshold is updated to 12.
[0131] In an embodiment of the present application, the first cache area stores input information with a query frequency greater than or equal to a first threshold and characteristic information of the information. Since the input information obtained when the user requests task processing is constantly changing and is always greater than or equal to the first threshold, the first threshold is adaptively adjusted according to the user's query frequency for information, so that the first threshold increases as the user's query frequency for certain information increases, thereby being able to store information with a higher user query frequency and characteristic information of the information in the first cache area, while avoiding saving some information with a lower user query frequency, avoiding waste of cache space, and thus efficiently utilizing cache space.
[0132] The following is an example of the process of adaptive system_prompt pre-processing cache provided by the embodiment of the present application.
[0133] Exemplarily, the task processing method may include the following steps:
[0134] Step 41: The task processing device determines whether the current system_prompt (ie, the first sub-information) exists in the cache sys_kv_map. If so, step 42a is executed; if not, step 42b is executed.
[0135] Step 42a: Obtain the cache result (ie, the first feature information) from sys_kv_map.
[0136] Step 42b: Calculate K and V of system_prompt.
[0137] Optionally, step 42b may be followed by steps 43 and 44.
[0138] Step 43: Get the request frequency n of system_prompt.
[0139] Step 44: Determine whether the request frequency n reaches the cache threshold N. If so, proceed to step 45; if not, end.
[0140] Step 45: Get the sum of the token lengths of all system_prompts in sys_kv_map (cache_seq_count).
[0141] Step 46: Determine whether cache_seq_count exceeds the limit MAX_SEQ_LIMIT. If so, execute steps 47a and 48; if not, execute step 47b.
[0142] Step 47a: Replace the record with the least frequent access in sys_kv_map and update cache_seq_count.
[0143] Step 48: Set the cache admission condition N to the minimum access frequency in sys_kv_map.
[0144] Step 47b: Store the KV cache in sys_kv_map and update cache_seq_count.
[0145] GPU cache resources are extremely valuable and very limited. Furthermore, as business models become increasingly complex, a mature large-scale model application may have hundreds of prompt inputs, which will continue to change with business iterations. Defining caches through configuration files has the following limitations: (1) Resource management is difficult. Configuring too much or too little system_prompt cache can lead to insufficient or wasted GPU cache resources. (2) It is impossible to achieve a global optimal solution. User requests change regularly every day, which may result in the system_prompt used in a certain period of time not being in the cache. Fixed caching strategies cannot achieve the most efficient utilization.
[0146] The embodiment of the present application provides a solution for adaptively inputting prompt pre-processing cache according to the frequency of user requests, which efficiently uses GPU cache resources and improves the response speed of the first word of large-model reasoning. Through this implementation, the interactive experience of users when using large-model applications can be greatly improved, greatly reducing the waiting time for responses. At the same time, due to the reduction of a large amount of repeated calculations, the efficiency of the service provider's GPU computing resources is also greatly improved, further promoting the popularization and development of artificial intelligence applications.
[0147] The above-mentioned method embodiments, or various possible implementation methods in each method embodiment, can be executed separately, or, under the premise that there is no contradiction, can also be executed in combination with each other. The specific implementation can be determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0148] The task processing method provided in the embodiment of the present application can be executed by a task processing device. In the embodiment of the present application, the task processing device provided in the embodiment of the present application is described by taking the task processing method executed by the task processing device as an example.
[0149] Figure 5 is a schematic diagram of a task processing device provided in an embodiment of the present application. As shown in Figure 5, the task processing device 500 may include an acquisition module 501 and a processing module 502, wherein: the acquisition module 501 is used to obtain the first prompt information of the first task, and split the first prompt information into first sub-information and second sub-information, the first sub-information indicates the business scenario of the first task, and the second sub-information indicates the intention of the first task; the acquisition module 501 is also used to obtain the first feature information of the first sub-information from the first cache area, the first cache area includes at least one feature information, and each feature information corresponds to a type of business scenario; the processing module 502 is used to obtain the first reasoning result of the first task through reasoning processing of the first feature information and the second feature information by the first AI model, and the second feature information is obtained by transforming the second sub-information.
[0150] In some embodiments of the present application, the at least one characteristic information is characteristic information of at least one first information, and each first information is used to indicate a type of business scenario; the acquisition module 501 is specifically used to: match the first sub-information with at least one first information in the first cache area; and determine the characteristic information in the at least one first information that matches the first sub-information as the first characteristic information.
[0151] In some embodiments of the present application, the task processing device 500 further includes: a storage module; the acquisition module 501 is further used to obtain the second prompt information of the second task before obtaining the first prompt information of the first task, and split the second prompt information of the second task into third sub-information and fourth sub-information, wherein the third sub-information indicates the business scenario of the second task, and the fourth sub-information indicates the intention of the second task; the processing module 502 is further used to transform the third sub-information according to the transformation matrix to obtain third feature information; the storage module is used to store the third sub-information and the third feature information in the first cache area when the query frequency of the third sub-information is greater than or equal to the first threshold.
[0152] In some embodiments of the present application, the above-mentioned processing module 502 is also used to delete the second information and the characteristic information of the second information in the above-mentioned first cache area before storing the above-mentioned third sub-information and third characteristic information in the first cache area, when the amount of data cached in the first cache area is greater than or equal to the second threshold; wherein the above-mentioned second information is: information with the minimum query frequency or the query frequency is less than the third threshold.
[0153] In some embodiments of the present application, the above-mentioned processing module 502 is also used to update the first threshold to the query frequency of the third information after storing the above-mentioned third sub-information and third characteristic information in the first cache area, and the above-mentioned third information is the information with the lowest query frequency in the first cache area.
[0154] The task processing device provided in the embodiment of the present application, when performing task processing, splits the prompt information into two parts: information indicating the business scenario and information indicating the schematic diagram. The task processing device then directly obtains the characteristic information of the information indicating the business scenario from the cache area, calculates the characteristic information of the partial information indicating the schematic diagram in the prompt information, and then performs inference processing based on the characteristic information of these two parts of information. Therefore, there is no need to transform the entire prompt information to obtain the characteristic information of the entire prompt information. Instead, only the partial information indicating the schematic diagram in the prompt information needs to be transformed to obtain the characteristic information of the partial information, thereby reducing the amount of calculation in the processing process and improving processing efficiency.
[0155] The task processing device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.
[0156] The task processing device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0157] The task processing device provided in the embodiment of the present application can implement each process implemented in the method embodiments of Figures 1 to 4. To avoid repetition, they will not be described here.
[0158] Optionally, as shown in Figure 6, an embodiment of the present application also provides an electronic device 600, including a processor 601 and a memory 602, and the memory 602 stores a program or instruction that can be run on the processor 601. When the program or instruction is executed by the processor 601, the various steps of the above-mentioned task processing method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0159] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0160] FIG7 is a schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.
[0161] The electronic device 100 includes but is not limited to components such as a radio frequency unit 101 , a network module 102 , an audio output unit 103 , an input unit 104 , a sensor 105 , a display unit 106 , a user input unit 107 , an interface unit 108 , a memory 109 , and a processor 110 .
[0162] Those skilled in the art will appreciate that the electronic device 100 may further include a power source (e.g., a battery) to power various components. The power source may be logically connected to the processor 110 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The electronic device structure shown in FIG7 does not limit the electronic device. The electronic device may include more or fewer components than shown, or may combine certain components or arrange the components differently, which will not be described in detail here.
[0163] Among them, the above-mentioned processor 110 is used to obtain the first prompt information of the first task, and split the above-mentioned first prompt information into first sub-information and second sub-information, the above-mentioned first sub-information indicates the business scenario of the above-mentioned first task, and the above-mentioned second sub-information indicates the intention of the above-mentioned first task; the above-mentioned processor 110 is also used to obtain the first feature information of the above-mentioned first sub-information from the first cache area, and the above-mentioned first cache area includes at least one feature information, and each feature information corresponds to a type of business scenario; the above-mentioned processor 110 is used to infer and process the above-mentioned first feature information and the second feature information through the first AI model to obtain the first inference result of the above-mentioned first task, and the above-mentioned second feature information is obtained by transforming the second sub-information.
[0164] In some embodiments of the present application, the at least one characteristic information is characteristic information of at least one first information, and each first information is used to indicate a type of business scenario; the processor 110 is specifically configured to:
[0165] The first sub-information is matched with at least one first information in the first buffer area; and characteristic information in the at least one first information that matches the first sub-information is determined as the first characteristic information.
[0166] In some embodiments of the present application, the processor 110 is further used to obtain the second prompt information of the second task before obtaining the first prompt information of the first task, and split the second prompt information of the second task into third sub-information and fourth sub-information, wherein the third sub-information indicates the business scenario of the second task, and the fourth sub-information indicates the intention of the second task; the processor 110 is further used to transform the third sub-information according to the transformation matrix to obtain third feature information; the memory 109 is used to store the third sub-information and the third feature information in the first cache area when the query frequency of the third sub-information is greater than or equal to the first threshold.
[0167] In some embodiments of the present application, the processor 110 is further used to delete the second information and the characteristic information of the second information in the first cache area before storing the third sub-information and the third characteristic information in the first cache area when the amount of data cached in the first cache area is greater than or equal to a second threshold; wherein the second information is: information with the minimum query frequency or a query frequency less than the third threshold.
[0168] In some embodiments of the present application, the processor 110 is further used to update the first threshold to the query frequency of the third information after storing the third sub-information and the third characteristic information in the first cache area, and the third information is the information with the lowest query frequency in the first cache area.
[0169] The electronic device provided in the embodiment of the present application, when performing task processing, splits the prompt information into two parts: information indicating a business scenario and information indicating a schematic diagram. The electronic device then directly obtains the characteristic information of the information indicating the business scenario from the cache area, calculates the characteristic information of the partial information indicating the schematic diagram in the prompt information, and then performs inference processing based on the characteristic information of the two parts of information. Therefore, there is no need to transform the entire prompt information to obtain the characteristic information of the entire prompt information. Instead, only the partial information indicating the schematic diagram in the prompt information needs to be transformed to obtain the characteristic information of the partial information, thereby reducing the amount of calculation in the processing process and improving processing efficiency.
[0170] It should be understood that in an embodiment of the present application, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042, and the graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes a touch panel 1071 and at least one of other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
[0171] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0172] Processor 110 may include one or more processing units. Optionally, processor 110 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 110.
[0173] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned task processing method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0174] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0175] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned task processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0176] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0177] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned task processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0178] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0179] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0180] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A task processing method, the method comprising: Obtaining first prompt information for a first task, and splitting the first prompt information into first sub-information and second sub-information, wherein the first sub-information indicates a business scenario of the first task, and the second sub-information indicates an intention of the first task; Acquire first feature information of the first sub-information from a first buffer area, where the first buffer area includes at least one feature information, each of which corresponds to a type of business scenario; The first feature information and the second feature information are inferred and processed by a first artificial intelligence (AI) model to obtain a first inference result of the first task, and the second feature information is obtained by transforming the second sub-information.
2. The method according to claim 1, wherein The at least one characteristic information is characteristic information of at least one first information, each first information is used to indicate a type of business scenario; The acquiring the first characteristic information of the first sub-information from the first buffer area includes: matching the first sub-information with at least one first information in the first buffer area; The characteristic information of the information matching the first sub-information in the at least one first information is determined as the first characteristic information.
3. The method according to claim 1, before obtaining the first prompt information of the first task, the method further comprises: Obtaining second prompt information for the second task, and splitting the second prompt information for the second task into third sub-information and fourth sub-information, wherein the third sub-information indicates a business scenario of the second task, and the fourth sub-information indicates an intention of the second task; Performing transformation processing on the third sub-information according to the transformation matrix to obtain third feature information; When the query frequency of the third sub-information is greater than or equal to a first threshold, the third sub-information and the third characteristic information are stored in the first cache area.
4. The method according to claim 3, before storing the third sub-information and the third feature information in the first cache area, the method further comprises: When the amount of cached information in the first buffer area is greater than or equal to a second threshold, deleting the second information and the characteristic information of the second information in the first buffer area; The second information is information indicating that the query frequency is the minimum or the query frequency is less than a third threshold.
5. The method according to claim 3 or 4, after storing the third sub-information and the third feature information in the first cache area, the method further comprises: The first threshold is updated to the query frequency of third information, where the third information is the information with the smallest query frequency in the first cache area.
6. A task processing device, comprising: Acquisition module and processing module, where: The acquisition module is configured to acquire first prompt information of a first task and split the first prompt information into first sub-information and second sub-information, wherein the first sub-information indicates a business scenario of the first task and the second sub-information indicates an intention of the first task; The acquisition module is further configured to acquire first feature information of the first sub-information from a first buffer area, where the first buffer area includes at least one feature information, each feature information corresponding to a type of business scenario; The processing module is used to perform inference processing on the first feature information and the second feature information through a first AI model to obtain a first inference result of the first task, and the second feature information is obtained by transforming the second sub-information.
7. The device according to claim 6, wherein The at least one characteristic information is characteristic information of at least one first information, each first information is used to indicate a type of business scenario; the acquisition module is specifically used to: matching the first sub-information with at least one first information in the first buffer area; The characteristic information in the at least one first information that matches the first sub-information is determined as the first characteristic information.
8. The apparatus according to claim 6, further comprising: Storage module; The acquisition module is further configured to acquire second prompt information for the second task before acquiring the first prompt information for the first task, and split the second prompt information for the second task into third sub-information and fourth sub-information, wherein the third sub-information indicates a business scenario for the second task, and the fourth sub-information indicates an intention for the second task; The processing module is further configured to perform transformation processing on the third sub-information according to the transformation matrix to obtain third feature information; The storage module is configured to store the third sub-information and the third characteristic information in the first cache area when the query frequency of the third sub-information is greater than or equal to a first threshold.
9. The apparatus according to claim 8, wherein the processing module is further configured to, before storing the third sub-information and the third characteristic information in the first cache area, delete the second information and the characteristic information of the second information in the first cache area if the amount of data cached in the first cache area is greater than or equal to a second threshold; in, The second information is information that the query frequency is the minimum or the query frequency is less than a third threshold.
10. In the device according to claim 8 or 9, the processing module is further used to update the first threshold to the query frequency of the third information after storing the third sub-information and the third characteristic information in the first cache area, and the third information is the information with the lowest query frequency in the first cache area.
11. An electronic device comprising a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
12. A readable storage medium storing a program or instruction, wherein the program or instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 5.
13. A computer program product, wherein the program product is stored in a storage medium and, when the program product is executed by at least one processor, implements the steps of the method according to any one of claims 1 to 5.
14. A chip comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to implement the steps of the method according to any one of claims 1 to 5 when executing a program or an instruction.
15. An electronic device, configured to perform the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Data processing method, server and electronic device
CN109063100A
Information processing method and device
CN111752982A
Task processing method and device
CN118152006A