Task processing method and automatic question answering method
By obtaining and distributing feature block sequences in large-model services, the problem of memory limitation in large-model services in long-context tasks is solved, achieving higher resource utilization and task processing flexibility.
Patent Information
- Application Number
- CN202410010610.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-08
AI Technical Summary
Due to the dynamic autoregression characteristics, large model services cannot determine the life cycle and sequence length of the task processing process in advance, resulting in extremely low flexibility and adaptability of the task processing process, especially in long context-length tasks, which are prone to exceed the GPU memory limit of the computing instance.
By obtaining the feature block sequence of target task data, when the current service unit is insufficient, the target service unit is queried from other service units, and the current and target service units are called to distribute the feature blocks to generate the final task processing result.
Improve resource utilization, support longer context-length task processing, enhance the flexibility and adaptability of task processing, and avoid performance fluctuations during data exchange or real-time migration.
Smart Images

Figure CN120277175A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a task processing method and an automatic question answering method. Background Art
[0002] With the development of computer technology, large models have begun to shine, showing extraordinary capabilities in language understanding, generation, interaction, and reasoning, and are widely used in processing fields such as dialogue, translation, and code generation. The rapid development of large models has gradually become the driving force for the growth of cloud-based large model services, which have now become an important part of promoting artificial intelligence applications.
[0003] However, due to the dynamic autoregressive characteristics of large model services, it is impossible to pre-determine the life cycle and sequence length of the large model task processing process. Therefore, the ability of large models to handle tasks with long context lengths is limited, resulting in extremely low flexibility and adaptability in the task processing process. Therefore, there is an urgent need for a task processing solution with high flexibility and adaptability. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a task processing method. One or more embodiments of this specification simultaneously relate to an automatic question answering method, a task processing device, an automatic question answering device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a task processing method is provided, including:
[0006] Obtain a sequence of feature blocks of target task data, where the sequence of feature blocks includes multiple feature blocks;
[0007] In the case of insufficient memory in the current service unit, query a target service unit from service units other than the current service unit, where the target service unit includes available memory for processing feature blocks;
[0008] Call the current service unit to process the first feature block among the multiple feature blocks to obtain a first processing result, and call the target service unit to process the second feature block among the multiple feature blocks to obtain a second processing result;
[0009] Generate a task processing result of the target task data according to the first processing result and the second processing result.
[0010] According to the second aspect of the embodiments of this specification, an automatic question answering method is provided, including:
[0011] Obtain a sequence of feature blocks of the question to be answered, where the sequence of feature blocks includes multiple feature blocks;
[0012] When the memory of the current service unit is insufficient, a target service unit is queried from service units other than the current service unit, where the target service unit includes available memory for processing feature blocks.
[0013] Call the current service unit to process the first feature block among multiple feature blocks to obtain a first processing result, and call the target service unit to process the second feature block among multiple feature blocks to obtain a second processing result.
[0014] Generate a reply result to the question to be answered according to the first processing result and the second processing result.
[0015] According to the third aspect of the embodiments of this specification, a task processing device is provided, including:
[0016] A first acquisition module, configured to acquire a sequence of feature blocks of target task data, where the sequence of feature blocks includes multiple feature blocks;
[0017] A first query module, configured to query a target service unit from service units other than the current service unit when the memory of the current service unit is insufficient, where the target service unit includes available memory for processing feature blocks;
[0018] A first processing module, configured to call the current service unit to process the first feature block among multiple feature blocks to obtain a first processing result, and call the target service unit to process the second feature block among multiple feature blocks to obtain a second processing result;
[0019] A first generation module, configured to generate a task processing result of the target task data according to the first processing result and the second processing result.
[0020] According to the fourth aspect of the embodiments of this specification, an automatic question answering device is provided, including:
[0021] A second acquisition module, configured to acquire a sequence of feature blocks of the question to be answered, where the sequence of feature blocks includes multiple feature blocks;
[0022] A second query module, configured to query a target service unit from service units other than the current service unit when the memory of the current service unit is insufficient, where the target service unit includes available memory for processing feature blocks;
[0023] A second processing module, configured to call the current service unit to process the first feature block among multiple feature blocks to obtain a first processing result, and call the target service unit to process the second feature block among multiple feature blocks to obtain a second processing result;
[0024] The second generating module is configured to generate a reply result of the question to be answered according to the first processing result and the second processing result.
[0025] According to a fifth aspect of an embodiment of this specification, a computing device is provided, including:
[0026] Memory and processor;
[0027] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in the first aspect or the second aspect are implemented.
[0028] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the method provided in the first aspect or the second aspect are implemented.
[0029] According to a seventh aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in the first aspect or the second aspect above.
[0030] A task processing method provided by an embodiment of the present specification includes: obtaining a feature block sequence of target task data, wherein the feature block sequence includes multiple feature blocks; when the current service unit has insufficient memory, querying a target service unit from a service unit other than the current service unit, wherein the target service unit includes available memory for processing feature blocks; calling the current service unit to process a first feature block among the multiple feature blocks to obtain a first processing result, and calling the target service unit to process a second feature block among the multiple feature blocks to obtain a second processing result; and generating a task processing result of the target task data based on the first processing result and the second processing result. By calling the target service unit for processing, the underutilized resources in the target service unit are fully utilized, and the resource utilization rate is improved. In addition, by dividing the multiple feature blocks into the first feature block and the second feature block for distributed processing, it is possible to support target task data with a longer context length, and the flexibility and adaptability of task processing are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is an architecture diagram of a task processing system provided by an embodiment of this specification;
[0032] Figure 2 is an architecture diagram of another task processing system provided by an embodiment of this specification;
[0033] Figure 3It is a flowchart of a task processing method provided by an embodiment of this specification;
[0034] Figure 4 It is a process flowchart of a task processing method provided by an embodiment of this specification;
[0035] Figure 5 It is a flowchart of an automatic question answering method provided by an embodiment of this specification;
[0036] Figure 6 It is a schematic diagram of the interface of an automatic question answering interface provided by an embodiment of this specification;
[0037] Figure 7 It is a schematic diagram of the structure of a task processing device provided by an embodiment of this specification;
[0038] Figure 8 It is a schematic diagram of the structure of an automatic question answering device provided by an embodiment of this specification;
[0039] Figure 9 It is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0040] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0041] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0042] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0043] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0044] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one quadrillion model parameters. A large model can also be called a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one billion parameters is produced. This kind of model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0045] When a large model is actually applied, only a small number of samples are needed to fine-tune the pre-trained model for application in different tasks. Large models can be widely applied in the fields of natural language processing (NLP), computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., and tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0046] First, the noun terms involved in one or more embodiments of this specification are explained.
[0047] Distributed attention mechanism: The Distributed Attention mechanism (DistAttention) is used to improve the efficiency and performance of the model when processing large-scale data.
[0048] Distributed Large Language Model Service System: The Distributed Key-Value Large Language Model (DistKVLLM) is a large language model service system based on a distributed architecture that can support a large number of user requests and handle high-concurrency requests.
[0049] KV Cache: The Key-Value Cache (KV Cache) is a distributed cache system used to store and manage data, improving data access speed and efficiency.
[0050] Cross-Data Center: Cross-data center means that the service system can transfer and process data between multiple data centers, improving the availability and reliability of the system.
[0051] Graphics Processing Unit: The Graphics Processing Unit (GPU) is used to accelerate the processing of computationally intensive tasks such as deep learning.
[0052] Core Processor Memory: Core processor memory refers to the memory of the Central Processing Unit (CPU), which is used to store and process data.
[0053] End-to-End Throughput: End-to-end throughput refers to the processing speed and efficiency of the service system from user requests to response results.
[0054] Context Length: Context length refers to the maximum text length that the model can process, which is used to improve the accuracy and reliability of the model.
[0055] Performance Improvement: Performance improvement refers to the performance improvement of the service system when processing large-scale data and high-concurrency requests, which can better meet the needs of users.
[0056] Large language models have driven the rapid growth of LLM services and become a key infrastructure for promoting the development of artificial intelligence applications. However, this development faces significant challenges due to the huge computational and data requirements. These services typically use multiple GPU cards to cooperate in completing LLM tasks. However, the dynamic nature of LLMs brings complex computational problems. The core of LLM services is the inherent process of autoregressive text generation, where the model generates one word (or token) at a time. Each newly generated token is appended to the existing text corpus, which constitutes the input for the internal calibration of the LLM. This iterative process continues until the final word or token is generated. Crucially, the memory and computational resources required for LLM services fluctuate continuously throughout the LLM service process, and neither their lifecycle nor sequence length can be known in advance.
[0057] The dynamic and iterative nature of autoregressive text generation makes it impossible to pre-plan resource allocation, which poses a substantial challenge when designing an efficient LLM service system on the cloud. Especially in long-context tasks, the continuously expanding KV cache may exceed the GPU memory limit in the computing instance. In such cases, immediate resource reallocation is required. This typically involves launching expensive real-time migrations to transfer tasks to more capable instances or pre-allocate additional GPUs to handle potential memory overloads. However, especially in tasks with normal context lengths, the latter may lead to inefficiencies and resource waste.
[0058] Currently, the above problems can generally be solved by facilitating data exchange between GPU and CPU memory. However, this method encounters several limitations. First, the memory exchange scope is limited to the GPU and CPU memory within a single node, so its ability to accommodate extremely long context lengths is restricted. Second, since this scheme exchanges the entire KV cache at the request level, it misses the opportunity for more adaptive and fine-grained scheduling in a distributed cloud environment. Finally, the computational interruption of the requests swapped out may cause fluctuations in the performance of the running tasks, which may fail to meet the requirements of the strict service level agreements crucial for cloud services.
[0059] To solve the above problems, the embodiments of this specification propose a task processing method that obtains a sequence of feature blocks of target task data, where the sequence of feature blocks includes multiple feature blocks; in the case of insufficient memory in the current service unit, a target service unit is queried from service units other than the current service unit, where the target service unit includes available memory for processing feature blocks; the current service unit is called to process the first feature block among the multiple feature blocks to obtain a first processing result, and the target service unit is called to process the second feature block among the multiple feature blocks to obtain a second processing result; according to the first processing result and the second processing result, a task processing result of the target task data is generated. By calling the target service unit for processing, the underutilized resources in the target service unit are fully utilized, improving resource utilization. Moreover, by dividing the multiple feature blocks into the first feature block and the second feature block for distributed processing, it is possible to support target task data with longer context lengths, improving the flexibility and adaptability of task processing.
[0060] In this specification, a task processing method is provided. This specification also relates to an automatic question-answering method, a task processing device, an automatic question-answering device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0061] See Figure 1 , Figure 1The architecture diagram of a task processing system provided by an embodiment of this specification is shown. The task processing system may include a client 100 and a server 200, and the server 200 includes multiple service units;
[0062] The client 100 is configured to send a sequence of feature blocks of target task data to the server 200, where the sequence of feature blocks includes multiple feature blocks;
[0063] The server 200 is configured to, when the memory of the current service unit is insufficient, query a target service unit from service units other than the current service unit, where the target service unit includes available memory for processing feature blocks; call the current service unit to process the first feature block among the multiple feature blocks to obtain a first processing result, and call the target service unit to process the second feature block among the multiple feature blocks to obtain a second processing result; generate a task processing result of the target task data according to the first processing result and the second processing result; send the task processing result to the client 100;
[0064] The client 100 is further configured to receive the task processing result sent by the server 200.
[0065] Applying the solution of the embodiment of this specification, by calling the target service unit for processing, the resources that are not fully utilized in the target service unit are fully utilized, improving the resource utilization rate. Moreover, by dividing the multiple feature blocks into the first feature block and the second feature block for distributed processing, it is thus possible to support target task data with a longer context length, improving the flexibility and adaptability of task processing.
[0066] See Figure 2 , Figure 2 shows the architecture diagram of another task processing system provided by an embodiment of this specification. The task processing system may include multiple clients 100 and a server 200, where the client 100 may include an end-side device and the server 200 may include a cloud-side device. Communication connections can be established between the multiple clients 100 through the server 200. In a task processing scenario, the server 200 is used to provide task processing services between the multiple clients 100, and the multiple clients 100 can respectively act as a sending end or a receiving end to achieve communication through the server 200.
[0067] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In a task processing scenario, it may be that the user publishes a data stream to the server 200 through the client 100, and the server 200 generates a task processing result according to the data stream and pushes the task processing result to other clients that have established communication.
[0068] Among them, a connection is established between the client 100 and the server 200 through a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The data transmitted by the client 100 may need to be processed such as encoded, transcoded, compressed, etc. before being published to the server 200.
[0069] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application program), or a cloud application, etc. The client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run, etc. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0070] The server 200 can include servers that provide various services, such as a server that provides communication services for multiple clients, or a server for background training that supports the models used on the client, or a server that processes the data sent by the client, etc. It should be noted that the server 200 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.
[0071] It should be noted that the task processing method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have a similar function as the server, so as to execute the task processing method provided in the embodiments of this specification. In other embodiments, the task processing method provided in the embodiments of this specification may also be jointly executed by the client and the server.
[0072] See Figure 3 , Figure 3 which shows a flowchart of a task processing method provided by an embodiment of this specification, specifically including the following steps:
[0073] Step 302: Obtain the feature block sequence of the target task data, where the feature block sequence includes multiple feature blocks.
[0074] In one or more embodiments of this specification, in order to improve the flexibility and adaptability of task processing, the feature block sequence of the target task data can be obtained, and distributed task processing can be performed based on the feature block sequence.
[0075] Specifically, the target task data refers to the data to be processed corresponding to the target task. The target task can be different tasks in different scenarios, such as different scenarios like the e-commerce scenario and the meeting scenario, and different tasks like the question-and-answer task, the retrieval task, the abstract extraction task, and so on. The data to be processed can be data of different modalities, such as the voice data to be processed, the text data to be processed, the image data to be processed, and so on.
[0076] It should be noted that in order to better perform distributed task processing, the KV cache can be divided into smaller units, called sub-blocks (rBlocks). A feature block refers to a sub-block that stores the task data features of the target task data.
[0077] In practical applications, there are various ways to obtain the feature block sequence of the target task data, which is specifically selected according to the actual situation, and this specification does not make any limitations in this regard. In a possible implementation manner of this specification, the feature block sequence of the target task data can be read from other data acquisition devices or databases.
[0078] In another possible implementation manner of this specification, the above-mentioned obtaining of the feature block sequence of the target task data may include the following steps:
[0079] Obtain the target task data;
[0080] Perform feature extraction on the target task data to obtain a task feature sequence, where the task feature sequence includes multiple task data features;
[0081] Construct a feature block sequence according to the feature block division condition and multiple task data features.
[0082] Specifically, the feature block division condition is used to divide multiple task data features into at least one feature block. The feature block division condition includes, but is not limited to, the number of task data features in the feature block and the number of feature blocks in the feature block sequence, which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard.
[0083] It should be noted that there are multiple ways to obtain the target task data, which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation manner of this specification, the target task data sent by the user can be received. In another possible implementation manner of this specification, the target task data can be read from other data acquisition devices or databases.
[0084] Furthermore, there are multiple ways to extract features from the target task data to obtain the task feature sequence, which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation manner of this specification, the target task data can be input into a deep learning model (such as a convolutional neural network, a recurrent neural network, etc.) for feature extraction to obtain the task feature sequence. In another possible implementation manner of this specification, a word embedding model (such as Word2Vec) can be used to extract features from the target task data to obtain the task feature sequence.
[0085] Exemplarily, assume that the task feature sequence includes 12 task data features, and the feature block division condition is that the number of task data features in the feature block is 4. Then, the 1st - 4th task data features can be stored in sub - block 1 to obtain feature block 1, the 5th - 8th task data features can be stored in sub - block 2 to obtain feature block 2, and the 9th - 12th task data features can be stored in sub - block 3 to obtain feature block 3, and a feature block sequence can be constructed based on feature block 1, feature block 2, and feature block 3.
[0086] Applying the solution of the embodiments of this specification, the target task data is obtained; features are extracted from the target task data to obtain a task feature sequence, where the task feature sequence includes multiple task data features; a feature block sequence is constructed according to the feature block division condition and the multiple task data features, effectively decomposing the KV cache into smaller and more manageable sub - blocks, thereby facilitating distributed memory management across data centers.
[0087] Step 304: In the case where the memory of the current service unit is insufficient, query the target service unit from the service units other than the current service unit, where the target service unit includes available memory for processing feature blocks.
[0088] In one or more embodiments of this specification, after obtaining the feature block sequence of the target task data, further, it is possible to determine whether the memory of the current service unit is sufficient. In the case where the memory of the current service unit is insufficient, a target service unit is queried from service units other than the current service unit.
[0089] Specifically, the current service unit is the service unit among multiple service units of the task processing system that is used to process the target task data. Therefore, the current service unit can be referred to as the local service unit or local instance for task processing. Service units other than the current service unit can be referred to as remote service units or remote instances for task processing. The target service unit includes available memory for processing feature blocks. Therefore, the target service unit can be understood as an example with excess capacity.
[0090] It should be noted that in the case where the memory of the current service unit is insufficient, it means that the current service unit cannot complete the processing of multiple feature blocks. At this time, available memory space can be borrowed from service units with excess capacity to jointly complete the processing of multiple feature blocks. That is, a target service unit is queried from service units other than the current service unit.
[0091] In an alternative embodiment of this specification, before querying a target service unit from service units other than the current service unit, it is possible to determine whether the memory of the current service unit is insufficient; if the memory of the current service unit is sufficient to process multiple feature blocks, the current service unit can be directly called to process the multiple feature blocks to obtain a processing result, and a task processing result of the target task data is generated according to the processing result; if the memory of the current service unit is insufficient to process multiple feature blocks, a target service unit can be queried from service units other than the current service unit.
[0092] In practical applications, there are various ways to determine whether the memory of the current service unit is insufficient, which are specifically selected according to the actual situation, and this specification does not make any limitations in this regard. In a possible implementation manner of this specification, the current service unit can be called to process multiple feature blocks, and the memory usage of the current service unit is monitored in real time. In the case where the processing of the feature blocks fails, it is determined that the memory of the current service unit is insufficient.
[0093] In another possible implementation manner of this specification, it is possible to determine whether the memory of the current service unit is insufficient according to the processing memory corresponding to multiple feature blocks. That is, before querying a target service unit from service units other than the current service unit in the case where the memory of the current service unit is insufficient, the following steps may further be included:
[0094] Obtain the current available memory of the current service unit and determine the processing memory corresponding to multiple feature blocks;
[0095] When the currently available memory is less than the processing memory, it is determined that the memory of the current service unit is insufficient.
[0096] Specifically, the currently available memory refers to the memory space that is not occupied and can process feature blocks. The currently available memory can also be understood as the free memory of the current service unit. The processing memory refers to the memory space required to process feature blocks.
[0097] In practical applications, there are various ways to obtain the currently available memory of the current service unit and determine the processing memory corresponding to multiple feature blocks. The specific selection is based on the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation manner of this specification, the currently available memory and the processing memory information sent by the user can be received. In another possible implementation manner of this specification, the currently available memory of the current service unit and the processing memory corresponding to multiple feature blocks can be read from the task manager.
[0098] It should be noted that after obtaining the currently available memory of the current service unit and determining the processing memory corresponding to multiple feature blocks, the memory sizes of the currently available memory and the processing memory can be compared. If the currently available memory is greater than or equal to the processing memory, it means that the memory of the current service unit is sufficient and the task processing can be independently completed. If the currently available memory is less than the processing memory, it means that the memory of the current service unit is insufficient and the task processing cannot be independently completed. It is necessary to borrow available memory space from service units with excess capacity and cooperate with other service units to complete the processing of multiple feature blocks.
[0099] Applying the solution of the embodiments of this specification, the currently available memory of the current service unit is obtained, and the processing memory corresponding to multiple feature blocks is determined; when the currently available memory is less than the processing memory, it is determined that the memory of the current service unit is insufficient. By determining whether the memory of the current service unit is insufficient according to the currently available memory and the processing memory, it is possible to avoid interruptions in processing due to insufficient memory of the service unit during the task processing, ensuring the continuity of task processing.
[0100] In practical applications, there are various ways to query and obtain the target service unit from service units other than the current service unit. The specific selection is based on the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation manner of this specification, the processing memory of multiple feature blocks can be evenly divided, and the target service unit can be screened out from multiple service units according to the evenly divided processing memory. The current service unit and the target service unit can evenly process multiple feature blocks.
[0101] In another possible implementation of this specification, a target service unit can be queried from service units other than the current service unit based on the memory difference between the currently available memory and the processing memory corresponding to multiple feature blocks, so as to make full use of the available memory of the current service unit. That is, the querying of the target service unit from service units other than the current service unit can include the following steps:
[0102] Determine memory requirement information based on the memory difference between the currently available memory of the current service unit and the processing memory corresponding to multiple feature blocks;
[0103] Screen out multiple candidate service units from multiple service units according to the unit memory usage table and the memory requirement information, where the candidate service units include available memory that meets the memory requirement information;
[0104] Screen out the target service unit from multiple candidate service units.
[0105] It should be noted that the memory requirement information is used to represent the size of the memory space that the current service unit needs to borrow. For example, if the currently available memory is a and the processing memory is b, then the memory requirement information is b - a. The unit memory usage table is used to track and record the memory usage of each service unit. The unit memory usage table includes the available memory information of the service unit and the memory interaction information between service units. The memory interaction information includes the service unit ID, the memory interaction size, and the memory interaction relationship. For example, the memory of service unit 1 is 80GB, the available memory is 30GB, service unit 1 borrows 11GB of memory from service unit 2, and lends 20GB of memory to service unit 3.
[0106] Exemplarily, assume that the unit memory usage table includes service unit 1, service unit 2, and service unit 3. The available memory of service unit 1 is 30GB, the available memory of service unit 2 is 3GB, the available memory of service unit 3 is 50GB, and the memory requirement information is 18GB. Then, according to the unit memory usage table and the memory requirement information, the multiple candidate service units screened out from multiple service units are service unit 1 and service unit 3.
[0107] In practical applications, the task processing system can be understood as a distributed large-scale language model service system. In the distributed large-scale language model service system, a management unit (gManager) can be set up, and a processing unit (rManager) can be set up in each service unit. Multiple service units share one management unit. The processing unit is used to virtualize the GPU and CPU memory in the service unit, and at the same time provide a unified application programming interface (API, Application Programming Interface) to serve local and remote memory operations, including allocating memory for feature blocks and releasing memory when processing is no longer required. The management unit is used to run as a global coordinator, maintain the global memory information of all service units, and ensure effective, scalable, and consistent resource management among distributed service units.
[0108] When the memory of the current service unit is insufficient, the rManager of the current service unit will attempt to borrow GPU or CPU memory space from neighboring service units, that is, initiate a query to the gManager and send memory requirement information to the gManager. After receiving the query, the gManager can query the unit memory usage table and provide the address ID of the candidate service unit, which represents the service unit with available memory currently. The rManager selects the target service unit from multiple candidate service units.
[0109] Applying the solution of the embodiments of this specification, determine the memory requirement information according to the memory difference between the current available memory of the current service unit and the processing memory corresponding to multiple feature blocks; screen out multiple candidate service units from multiple service units according to the unit memory usage table and the memory requirement information, where the candidate service units include the available memory that meets the memory requirement information; screen out the target service unit from multiple candidate service units, thereby avoiding determining the service unit with insufficient available memory as the target service unit, preventing service unit overload, and ensuring the availability of the target service unit.
[0110] In practical applications, there are various ways to screen out the target service unit from multiple candidate service units, which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In a possible implementation manner of this specification, the target service unit can be randomly determined from multiple candidate service units.
[0111] In another possible implementation manner of this specification, the target service unit can be screened out from multiple candidate service units following the principles of locality and availability, that is, the above-mentioned screening out of the target service unit from multiple candidate service units may include the following steps:
[0112] Send a memory borrowing request to the candidate service unit according to the priorities of multiple candidate service units;
[0113] In the case of receiving the borrowing confirmation information returned by the first candidate service unit, determine the first candidate service unit as the target service unit, where the first candidate service unit is any one of multiple candidate service units.
[0114] It should be noted that the priorities of multiple candidate service units can be determined based on communication costs and / or available memory sizes. The multiple candidate service units can be sorted according to their priorities to determine the request sending queue. Further, based on the request sending queue, a memory borrowing request is sent to the candidate service unit with a higher priority. If the candidate service unit with a higher priority returns the borrowing confirmation information, the candidate service unit can be determined as the target service unit; if the candidate service unit with a higher priority returns the borrowing rejection information, a memory borrowing request can be sent to the next candidate service unit with a lower priority until the target service unit is found.
[0115] In practical applications, if a service unit receives memory borrowing requests sent by multiple remote service units simultaneously, it can borrow memory for the remote service units based on the first-come, first-served policy. If a certain candidate service unit returns the borrowing rejection information, at this time, forwarding the memory borrowing request to this candidate service unit can be paused until the candidate service unit has more available memory.
[0116] Applying the solution of the embodiments of this specification, according to the priorities of multiple candidate service units, a memory borrowing request is sent to the candidate service unit; in the case of receiving the borrowing confirmation information returned by the first candidate service unit, the first candidate service unit is determined as the target service unit. This ensures the balanced and orderly allocation of memory resources, realizes screening out the target service unit from multiple candidate service units based on relatively low communication costs and high available memory space, and alleviates potential task processing bottlenecks in the system.
[0117] In an optional embodiment of this specification, before screening out multiple candidate service units from multiple service units according to the unit memory usage table and memory requirement information, the following steps may further be included:
[0118] Obtain the heartbeat information of multiple service units, where the heartbeat information includes the available memory information and memory interaction information of the service units;
[0119] Construct a unit memory usage table according to the available memory information and memory interaction information of multiple service units.
[0120] Specifically, the available memory information includes but is not limited to available memory addresses and available memory sizes. The memory interaction information includes but is not limited to memory interaction sizes and memory interaction relationships. Memory interaction includes but is not limited to borrowing and lending between service unit memories.
[0121] In practical applications, rManager can periodically send the heartbeat information of service units to gManager, and gManager maintains a unit memory usage table for summarizing the memory space usage of service units, so that gManager does not need to carefully track each memory allocation or release operation in all service units.
[0122] Applying the solution of the embodiments of this specification, the heartbeat information of multiple service units is obtained, where the heartbeat information includes the available memory information and memory interaction information of the service units; according to the available memory information and memory interaction information of the multiple service units, a unit memory usage table is constructed, thereby simplifying the memory space borrowing process, reducing processing latency and achieving performance improvement of the overall system.
[0123] In practical applications, since each service unit in the system can act as both a creditor and a debtor of memory and interact with other service units for memory as needed. For example, a service unit processing a request with a long context may grow continuously and thus needs to borrow space from a remote service unit. Conversely, a service unit with short-lived requests will release memory space faster and then can lend the memory space to other service units or allocate it for new requests. This dynamicity can lead to deterioration of data locality: since service units often access data stored in remote memory locations, the system will incur significant performance losses, such as increased latency and reduced end-to-end throughput. Therefore, in the embodiments of this specification, a fragmented memory management strategy is proposed, which aims to offset the memory interaction relationship by strategically recalling the memory space it has lent out and exchanging local feature blocks based on the memory interaction relationship, that is, after obtaining the heartbeat information of multiple service units, the following steps may further be included:
[0124] Parse the memory interaction information to determine the memory interaction relationship between multiple service units;
[0125] Taking service units as nodes and memory interaction relationships as edges, construct a memory interaction graph;
[0126] Adjust the feature blocks corresponding to the service units according to the memory interaction graph.
[0127] Specifically, the memory interaction relationship refers to the relationship of memory borrowing or lending between multiple service units. The memory interaction graph is a directed graph, that is, the edges in the memory interaction graph are directed edges representing memory interaction from one service unit to another service unit.
[0128] In practical applications, there are various ways to analyze memory interaction information and determine the memory interaction relationships between multiple service units. Specific selection is made according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation of this specification, the memory interaction relationships can be parsed from the memory logs of multiple service units. In another possible implementation of this specification, resource management and monitoring tools can be used to parse out the memory interaction relationships from the memory interaction information.
[0129] It should be noted that there are various ways to adjust the feature blocks corresponding to service units according to the memory interaction graph. Specific selection is made according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation of this specification, the feature blocks corresponding to service units can be directly adjusted according to the memory interaction graph. In another possible implementation of this specification, to avoid adjustment failure caused by the continuous memory interaction of the service unit during the adjustment process, the service unit can be frozen first, and then the feature blocks corresponding to the service unit can be adjusted according to the memory interaction graph.
[0130] Applying the solution of the embodiments of this specification, analyze memory interaction information to determine the memory interaction relationships between multiple service units; construct a memory interaction graph with service units as nodes and memory interaction relationships as edges; adjust the feature blocks corresponding to service units according to the memory interaction graph. By adjusting the feature blocks corresponding to each service unit, the overall system efficiency and memory utilization rate are improved.
[0131] In an optional embodiment of this specification, the above adjustment of the feature blocks corresponding to service units according to the memory interaction graph may include the following steps:
[0132] Screen out the service units to be adjusted from the memory interaction graph, where the service units to be adjusted form a circular relationship structure in the memory interaction graph;
[0133] Adjust the corresponding feature blocks of the service units to be adjusted according to the memory interaction information of the service units to be adjusted.
[0134] It should be noted that when screening out the service units to be adjusted from the memory interaction graph, a node can be randomly selected in the memory interaction graph, and at least one circular relationship structure can be found by traversing the memory interaction graph. This circular relationship structure indicates that the involved nodes can offset the interactive memory between each other.
[0135] Exemplarily, assume that service unit 1 borrows service unit 2 to process feature block 1, service unit 2 borrows service unit 3 to process feature block 2, and service unit 3 borrows service unit 1 to process feature block 3. At this time, a circular relationship structure is formed among these three service units. At this time, feature block 1 can be returned and stored in service unit 1, feature block 2 can be returned and stored in service unit 2, and feature block 3 can be returned and stored in service unit 3 to offset the interactive memory among them.
[0136] Applying the solution of the embodiments of this specification, the service units to be adjusted are screened out from the memory interaction graph, where the service units to be adjusted form a circular relationship structure in the memory interaction graph; according to the memory interaction information of the service units to be adjusted, the corresponding feature blocks of the service units to be adjusted are adjusted, reducing the need for memory access to remote service units, thereby improving the locality of data and ultimately leading to a significant improvement in system performance.
[0137] Step 306: Invoke the current service unit to process the first feature block among multiple feature blocks to obtain a first processing result, and invoke the target service unit to process the second feature block among multiple feature blocks to obtain a second processing result.
[0138] In one or more embodiments of this specification, a sequence of feature blocks of target task data is obtained; when the memory of the current service unit is insufficient, after querying the target service unit from service units other than the current service unit, further, the current service unit can be invoked to process the first feature block among multiple feature blocks to obtain a first processing result, and the target service unit can be invoked to process the second feature block among multiple feature blocks to obtain a second processing result.
[0139] Specifically, when invoking a service unit to process a feature block, the processing result can be generated based on a distributed attention mechanism. That is, when invoking the current service unit to process the first feature block among multiple feature blocks, it can be to invoke the current service unit to perform attention calculation on the first feature block among multiple feature blocks to obtain a first attention result. When invoking the target service unit to process the second feature block among multiple feature blocks, it can be to invoke the target service unit to perform attention calculation on the second feature block among multiple feature blocks to obtain a second attention result. Attention calculation (Attention Mechanism) is used to enable the model to dynamically allocate attention weights when processing input data. The attention calculation process generally includes the calculation of key vectors (Keys), value vectors (Values), and query vectors (Queries); the calculation of attention weights, where the weights reflect the correlation between the query vector and each key vector, that is, the importance of each part of the input sequence for the current task; the aggregation of attention weighted values: applying the attention weights to the corresponding value vectors; the generation of attention results.
[0140] In an alternative embodiment of this specification, since the current service unit has insufficient memory, the current service unit and the target service unit can cooperate to process multiple feature blocks. At this time, the original processing process can be divided into two micro-processing processes based on the number of service units participating in the attention processing, so as to complete the processing distributively. When each micro-processing is performed, it is necessary to determine its corresponding feature sequence. That is, before the above-mentioned steps of calling the current service unit to process the first feature block among multiple feature blocks to obtain a first processing result and calling the target service unit to process the second feature block among multiple feature blocks to obtain a second processing result, the following steps may further be included:
[0141] Obtain the current available memory of the current service unit;
[0142] Divide the multiple feature blocks according to the current available memory to determine the first feature block and the second feature block, where the current available memory is used to process the first feature block.
[0143] Exemplarily, assume that there are 2000 multiple feature blocks. The current available memory of the current service unit is 30GB, and 30GB can be used to process 1000 feature blocks. At this time, the multiple feature blocks can be divided according to the current available memory, and the first 1000 feature blocks can be determined as the first feature block, and the last 1000 feature blocks can be determined as the second feature block.
[0144] Applying the solution of the embodiment of this specification to obtain the current available memory of the current service unit; dividing the multiple feature blocks according to the current available memory to determine the first feature block and the second feature block ensures the accurate progress of the distributed processing.
[0145] In an alternative embodiment of this specification, the above-mentioned steps of calling the target service unit to process the second feature block among multiple feature blocks to obtain a second processing result may include the following steps:
[0146] Look up the target memory address in the address mapping table of the target service unit according to the attribute information of the second feature block;
[0147] Store the second feature block into the target service unit according to the target memory address, and call the target service unit to process the second feature block.
[0148] Specifically, the attribute information of the second feature block can be understood as the metadata of the second feature block, and the metadata is used to provide the logical sub-block information of the relevant sub-blocks. The address mapping table includes the mapping relationship between the logical sub-block information and the physical sub-block information. The memory address can be understood as the physical address.
[0149] It should be noted that the implementation method of calling the current service unit to process the first feature block among multiple feature blocks is the same as the implementation method of calling the target service unit to process the second feature block among multiple feature blocks to obtain the second processing result, and the embodiments of this specification will not elaborate on it.
[0150] See Figure 4 , Figure 4 shows the processing procedure flowchart of a task processing method provided by an embodiment of this specification, which specifically includes: during the task processing, the processing units in the local service unit and the remote service unit are used to divide the CPU and multiple GPUs into physical sub-blocks of a fixed size, effectively virtualizing the global memory. The processing units in multiple service units share a management unit. Moreover, the processing unit can also allocate memory for the feature block and release the memory when the feature block is no longer needed to be processed. The processing unit is used to maintain an address mapping table to map the logical sub-block to the corresponding physical sub-block in the global memory. As Figure 4 shown, the logical sub-block information includes the sub-block identifier (ID, Identity Document) and the service unit identifier. The sub-block identifier and the service unit identifier are used to indicate whether the KV cache of the feature block is stored in the local service unit or the remote service unit. The physical sub-block information includes the device identifier and the physical address identifier. The device identifier and the physical address identifier are used to indicate whether the physical memory address of the feature block is on the CPU side or one of the multiple GPUs.
[0151] In practical applications, for any service unit, the feature block can be processed through the following formula (1):
[0152]
[0153] where MA ij represents the processing result; exp represents the exponential function; Q represents the query vector; K represents the key vector; T represents the matrix transpose; i represents the i-th token in the feature block corresponding to the service unit; j represents the j-th key vector segmentation block of the feature block; represents the inner product result obtained by taking the inner product of the query vector of the i-th token and the j-th key vector segmentation block; represents the maximum value of the inner product result.
[0154] Applying the solution of the embodiments of this specification, according to the attribute information of the second feature block, the target memory address is searched from the address mapping table of the target service unit; according to the target memory address, the second feature block is stored in the target service unit, and the target service unit is called to process the second feature block, realizing distributed processing, so as to be able to support the target task data with a longer context length and improve the flexibility and adaptability of task processing.
[0155] Step 308: Generate a task processing result for the target task data based on the first processing result and the second processing result.
[0156] In one or more embodiments of this specification, a feature block sequence of the target task data is obtained; when the memory of the current service unit is insufficient, a target service unit is queried from service units other than the current service unit; after calling the current service unit to process the first feature block among multiple feature blocks to obtain a first processing result, and calling the target service unit to process the second feature block among multiple feature blocks to obtain a second processing result, further, a task processing result of the target task data can be generated based on the first processing result and the second processing result.
[0157] Specifically, the task processing result is related to the task corresponding to the target task data. For example, if the task is a question-and-answer task, the task processing result is a reply result; if the task is a retrieval task, the task processing result is a retrieval result.
[0158] Applying the solution of the embodiments of this specification, by calling the target service unit for processing, the resources that are not fully utilized in the target service unit are fully utilized, improving the resource utilization rate. And by dividing multiple feature blocks into a first feature block and a second feature block for distributed processing, it is possible to support target task data with a longer context length, improving the flexibility and adaptability of task processing. At the same time, performance fluctuations related to data exchange or real-time migration are avoided.
[0159] In practical applications, there are various ways to generate a task processing result for the target task data based on the first processing result and the second processing result, which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation manner of this specification, a first task processing result can be generated based on the first processing result, a second task processing result can be generated according to the second processing result, and a task processing result of the target task data can be generated based on the first task processing result and the second task processing result.
[0160] In another possible implementation manner of this specification, the above-mentioned generating a task processing result for the target task data based on the first processing result and the second processing result may include the following steps:
[0161] Aggregate the first processing result and the second processing result to obtain an aggregated processing result;
[0162] Generate a task processing result of the target task data according to the aggregated processing result.
[0163] In practical applications, aggregation refers to aggregating the results of distributed processing. When aggregating, the first processing result and the second processing result can be scaled and added to reduce, so as to obtain the aggregated processing result. Specifically, the first processing result and the second processing result can be aggregated through the following formulas (2), (3), and (4) to obtain the aggregated processing result:
[0164]
[0165]
[0166]
[0167] Among them, Attention(Q, K, V) represents the aggregated processing result; Reduce represents the aggregation operation; Scale represents the scaling operation; B kv represents the number of pairs of chunks into which the key vector and the value vector are sliced; max i represents the maximum value of the attention weights corresponding to the i-th query vector; represents the inner product of the query vector of the i-th token and the first chunk of the key vector; represents the inner product of the query vector of the i-th token and the B-th chunk of the key-value vector; sum i represents the denominator of the softmax calculation corresponding to the query vector of the i-th token.
[0168] Furthermore, after obtaining the aggregated processing result, the context vector can be obtained by weighted summation according to the aggregated processing result, and the task processing result of the target task data can be decoded from the context vector.
[0169] Applying the solution of the embodiments of this specification to aggregate the first processing result and the second processing result to obtain the aggregated processing result; according to the aggregated processing result, generate the task processing result of the target task data. Through distributed processing, it is thus possible to support target task data with a longer context length, improving the flexibility and adaptability of task processing.
[0170] The following combines the attached Figure 5 , taking the application of the task processing method provided in this specification in the intelligent question-answering scenario as an example, to further illustrate the task processing method. Among them, Figure 5 shows the flowchart of an automatic question-answering method provided by an embodiment of this specification, which specifically includes the following steps:
[0171] Step 502: Obtain the feature block sequence of the question to be answered, where the feature block sequence includes multiple feature blocks.
[0172] Step 504: When the memory of the current service unit is insufficient, query the target service unit from service units other than the current service unit, where the target service unit includes available memory for processing feature blocks.
[0173] Step 506: Invoke the current service unit to process the first feature block among multiple feature blocks to obtain a first processing result, and invoke the target service unit to process the second feature block among multiple feature blocks to obtain a second processing result.
[0174] Step 508: Generate a reply result to the question to be answered based on the first processing result and the second processing result.
[0175] It should be noted that the implementation manners of steps 502 to 508 can refer to the implementation manners of the above steps 302 to 308, and the embodiments of this specification will not be elaborated herein.
[0176] Applying the solution of the embodiments of this specification, by invoking the target service unit for processing, the resources that are not fully utilized in the target service unit are fully utilized, improving the resource utilization rate. Moreover, by dividing multiple feature blocks into the first feature block and the second feature block for distributed processing, it is possible to support questions to be answered with a longer context length, improving the flexibility and adaptability of automatic question answering.
[0177] See Figure 6 , Figure 6 shows a schematic diagram of an interface of an automatic question answering interface provided by an embodiment of this specification. The task processing interface is divided into a request input interface and a result display interface. The request input interface includes a request input box, an "OK" control, and a "Cancel" control. The result display interface includes a result display box.
[0178] The user inputs an automatic question answering request carrying a sequence of feature blocks of the question to be answered through the request input box displayed on the client, where the sequence of feature blocks includes multiple feature blocks; clicks on the "OK" control, and the server receives the sequence of feature blocks sent by the client. When the memory of the current service unit is insufficient, query the target service unit from service units other than the current service unit, where the target service unit includes available memory for processing feature blocks; invoke the current service unit to process the first feature block among multiple feature blocks to obtain a first processing result, and invoke the target service unit to process the second feature block among multiple feature blocks to obtain a second processing result; generate a reply result to the question to be answered based on the first processing result and the second processing result; and send the reply result to the client. The client displays the reply result in the result display box.
[0179] In practical applications, the ways for users to operate on a control include any one of clicking, double-clicking, touching, mouse hovering, swiping, long pressing, voice control, or shaking, etc., which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations on this.
[0180] Corresponding to the above embodiments of the task processing method, this specification also provides embodiments of a task processing device. Figure 7 The structural schematic diagram of a task processing device provided by an embodiment of this specification is shown. As Figure 7 shown, the device includes:
[0181] A first acquisition module 702, configured to acquire a sequence of feature blocks of target task data, where the sequence of feature blocks includes a plurality of feature blocks;
[0182] A first query module 704, configured to, when the current service unit has insufficient memory, query a target service unit from service units other than the current service unit, where the target service unit includes available memory for processing feature blocks;
[0183] A first processing module 706, configured to call the current service unit to process the first feature block among the plurality of feature blocks to obtain a first processing result, and call the target service unit to process the second feature block among the plurality of feature blocks to obtain a second processing result;
[0184] A first generation module 708, configured to generate a task processing result of the target task data according to the first processing result and the second processing result.
[0185] Optionally, the first query module 704 is further configured to determine memory requirement information according to the memory difference between the current available memory of the current service unit and the processing memory corresponding to the plurality of feature blocks; screen a plurality of candidate service units from the plurality of service units according to the unit memory usage table and the memory requirement information, where the candidate service units include available memory that meets the memory requirement information; screen out the target service unit from the plurality of candidate service units.
[0186] Optionally, the first query module 704 is further configured to send a memory borrowing request to the candidate service units according to the priorities of the plurality of candidate service units; and when receiving borrowing confirmation information returned by the first candidate service unit, determine the first candidate service unit as the target service unit, where the first candidate service unit is any one of the plurality of candidate service units.
[0187] Optionally, the device further includes: a construction module configured to obtain heartbeat information of a plurality of service units, where the heartbeat information includes available memory information and memory interaction information of the service units; and construct a unit memory usage table according to the available memory information and memory interaction information of the plurality of service units.
[0188] Optionally, the device further includes: an adjustment module configured to parse the memory interaction information, determine the memory interaction relationship between the plurality of service units; construct a memory interaction graph with the service units as nodes and the memory interaction relationship as edges; and adjust the feature blocks corresponding to the service units according to the memory interaction graph.
[0189] Optionally, the adjustment module is further configured to screen out service units to be adjusted from the memory interaction graph, where the service units to be adjusted form a circular relationship structure in the memory interaction graph; and adjust the feature blocks corresponding to the service units to be adjusted according to the memory interaction information of the service units to be adjusted.
[0190] Optionally, the first processing module 706 is further configured to look up a target memory address in the address mapping table of the target service unit according to the attribute information of the second feature block; store the second feature block into the target service unit according to the target memory address; and call the target service unit to process the second feature block.
[0191] Optionally, the first generation module 708 is further configured to aggregate the first processing result and the second processing result to obtain an aggregated processing result; and generate a task processing result of the target task data according to the aggregated processing result.
[0192] Optionally, the first acquisition module 702 is further configured to obtain target task data; perform feature extraction on the target task data to obtain a task feature sequence, where the task feature sequence includes a plurality of task data features; and construct a feature block sequence according to the feature block division condition and the plurality of task data features.
[0193] Optionally, the device further includes: a division module configured to obtain the current available memory of the current service unit; divide a plurality of feature blocks according to the current available memory to determine a first feature block and a second feature block, where the current available memory is used to process the first feature block.
[0194] By applying the solution of the embodiments of this specification and processing by calling the target service unit, the resources that are not fully utilized in the target service unit are fully utilized, improving the resource utilization rate. Moreover, by dividing a plurality of feature blocks into a first feature block and a second feature block for distributed processing, it is possible to support target task data with a longer context length, improving the flexibility and adaptability of task processing.
[0195] The above is a schematic solution of a task processing device according to this embodiment. It should be noted that the technical solution of this task processing device and the technical solution of the above task processing method belong to the same concept. For the details not described in detail in the technical solution of the task processing device, reference can be made to the description of the technical solution of the above task processing method.
[0196] Corresponding to the above embodiment of the automatic question answering method, this specification also provides an embodiment of an automatic question answering device. Figure 8 The structure diagram of an automatic question answering device provided by an embodiment of this specification is shown. As Figure 8 shown, the device includes:
[0197] A second acquisition module 802, configured to acquire a sequence of feature blocks of a question to be answered, where the sequence of feature blocks includes a plurality of feature blocks;
[0198] A second query module 804, configured to query a target service unit from service units other than the current service unit when the memory of the current service unit is insufficient, where the target service unit includes available memory for processing feature blocks;
[0199] A second processing module 806, configured to call the current service unit to process a first feature block among the plurality of feature blocks to obtain a first processing result, and call the target service unit to process a second feature block among the plurality of feature blocks to obtain a second processing result;
[0200] A second generation module 808, configured to generate a reply result to the question to be answered according to the first processing result and the second processing result.
[0201] Applying the solution of the embodiment of this specification, by calling the target service unit for processing, the resources not fully utilized in the target service unit are fully utilized, improving the resource utilization rate. Moreover, by dividing the plurality of feature blocks into a first feature block and a second feature block for distributed processing, it is possible to support questions to be answered with a longer context length, improving the flexibility and adaptability of automatic question answering.
[0202] The above is a schematic solution of an automatic question answering device according to this embodiment. It should be noted that the technical solution of this automatic question answering device and the technical solution of the above automatic question answering method belong to the same concept. For the details not described in detail in the technical solution of the automatic question answering device, reference can be made to the description of the technical solution of the above automatic question answering method.
[0203] Figure 9The block diagram of a computing device provided by an embodiment of this specification is shown. The components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.
[0204] The computing device 900 further includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interfaces (e.g., a Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Networks (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (WiMAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0205] In an embodiment of this specification, the above components of the computing device 900 and Figure 9 other components not shown in Figure 9 may also be connected to each other, for example, via a bus. It should be understood that
[0206] the block diagram of the computing device shown is only for illustrative purposes and is not a limitation on the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0206] The computing device 900 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or a Personal Computer (PC). The computing device 900 can also be a mobile or stationary server.
[0207] Among them, the processor 920 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above-mentioned task processing method or automatic question-answering method are implemented.
[0208] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above-mentioned task processing method and automatic question-answering method belong to the same concept. For the detailed content not described in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above-mentioned task processing method or automatic question-answering method.
[0209] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the above-mentioned task processing method or automatic question-answering method are implemented.
[0210] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above-mentioned task processing method and automatic question-answering method belong to the same concept. For the detailed content not described in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above-mentioned task processing method or automatic question-answering method.
[0211] An embodiment of this specification also provides a computer program. Among them, when the computer program is executed on a computer, the computer is made to execute the steps of the above-mentioned task processing method or automatic question-answering method.
[0212] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solutions of the above-mentioned task processing method and automatic question-answering method belong to the same concept. For the detailed content not described in the technical solution of the computer program, reference can be made to the descriptions of the technical solutions of the above-mentioned task processing method or automatic question-answering method.
[0213] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0214] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0215] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0216] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0217] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. The embodiments selected and specifically described in this specification are to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A task processing method, comprising: Obtaining a sequence of feature blocks of target task data, where the sequence of feature blocks includes a plurality of feature blocks; When the memory of the current service unit is insufficient, querying a target service unit from service units other than the current service unit, where the target service unit includes available memory for processing the feature blocks; Invoking the current service unit to process a first feature block among the plurality of feature blocks to obtain a first processing result, and invoking the target service unit to process a second feature block among the plurality of feature blocks to obtain a second processing result; Generating a task processing result of the target task data according to the first processing result and the second processing result.
2. The method according to claim 1, where the querying a target service unit from service units other than the current service unit includes: Determining memory requirement information according to the difference between the current available memory of the current service unit and the processing memory corresponding to the plurality of feature blocks; Filtering out a plurality of candidate service units from a plurality of service units according to a unit memory usage table and the memory requirement information, where the candidate service units include available memory that meets the memory requirement information; Filtering out a target service unit from the plurality of candidate service units.
3. The method according to claim 2, where the filtering out a target service unit from the plurality of candidate service units includes: Sending a memory borrowing request to the candidate service units according to the priorities of the plurality of candidate service units; When receiving borrowing confirmation information returned by a first candidate service unit, determining the first candidate service unit as the target service unit, where the first candidate service unit is any one of the plurality of candidate service units.
4. The method according to claim 2, before filtering out a plurality of candidate service units from a plurality of service units according to the unit memory usage table and the memory requirement information, further comprising: Obtaining heartbeat information of a plurality of service units, where the heartbeat information includes available memory information and memory interaction information of the service units; Constructing a unit memory usage table according to the available memory information and the memory interaction information of the plurality of service units.
5. The method according to claim 4, after obtaining the heartbeat information of a plurality of service units, further comprising: Analyzing the memory interaction information to determine the memory interaction relationship between the plurality of service units; Constructing a memory interaction graph with the service units as nodes and the memory interaction relationship as edges; Adjusting the feature blocks corresponding to the service units according to the memory interaction graph.
6. The method according to claim 5, where the adjusting the feature blocks corresponding to the service units according to the memory interaction graph includes: Filtering out service units to be adjusted from the memory interaction graph, where the service units to be adjusted form a circular relationship structure in the memory interaction graph; Adjusting the feature blocks corresponding to the service units to be adjusted according to the memory interaction information of the service units to be adjusted.
7. The method according to claim 1, wherein the invoking the target service unit to process the second feature block among the multiple feature blocks to obtain a second processing result includes: Looking up a target memory address from an address mapping table of the target service unit according to the attribute information of the second feature block; Storing the second feature block into the target service unit according to the target memory address, and invoking the target service unit to process the second feature block.
8. The method according to claim 1, wherein the generating a task processing result of the target task data according to the first processing result and the second processing result includes: Aggregating the first processing result and the second processing result to obtain an aggregated processing result; Generating a task processing result of the target task data according to the aggregated processing result.
9. The method according to claim 1, wherein the obtaining a sequence of feature blocks of the target task data includes: Obtaining the target task data; Performing feature extraction on the target task data to obtain a task feature sequence, where the task feature sequence includes multiple task data features; Constructing a sequence of feature blocks according to a feature block division condition and the multiple task data features.
10. Before the method according to claim 1, where the current service unit is invoked to process the first feature block among the multiple feature blocks to obtain a first processing result, and the target service unit is invoked to process the second feature block among the multiple feature blocks to obtain a second processing result, further includes: Obtaining the current available memory of the current service unit; Dividing the multiple feature blocks according to the current available memory to determine a first feature block and a second feature block, where the current available memory is used to process the first feature block.
11. An automatic question answering method, including: Obtaining a sequence of feature blocks of a question to be answered, where the sequence of feature blocks includes multiple feature blocks; When the memory of the current service unit is insufficient, querying a target service unit from service units other than the current service unit, where the target service unit includes available memory for processing the feature blocks; Invoking the current service unit to process the first feature block among the multiple feature blocks to obtain a first processing result, and invoking the target service unit to process the second feature block among the multiple feature blocks to obtain a second processing result; Generating a reply result to the question to be answered according to the first processing result and the second processing result.
12. A computing device, including: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10 or claim 11 are implemented.
13. A computer-readable storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 or claim 11 are implemented.