Task processing method and automatic question answering method

By obtaining feature block sequences in large-model services and distribute them when memory is insufficient, the problem of resource waste and performance fluctuations in large-model services in long-context tasks is solved, and more efficient task processing is achieved.

WO2025146635A1PCT designated stage expired Publication Date: 2025-07-10CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/050023
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2025-01-02
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Due to the dynamic autoregression characteristics, large model services cannot determine the life cycle and sequence length of the task processing process in advance, resulting in extremely low flexibility and adaptability of the task processing process, especially in long-context tasks, which can easily lead to resource waste and performance fluctuations.

Method used

By obtaining the feature block sequence of target task data, when the current service unit is insufficient, the target service unit is queried from other service units, and the current and target service units are called to distribute the feature blocks to generate task processing results.

Benefits of technology

Improve resource utilization, support longer context-length task processing, enhance the flexibility and adaptability of task processing, and avoid performance fluctuations during data exchange or real-time migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050023_10072025_PF_FP_ABST
    Figure IB2025050023_10072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a task processing method and an automatic question answering method. The task processing method comprises: acquiring a feature block sequence of target task data, wherein the feature block sequence comprises a plurality of feature blocks; in the case of insufficient memory in the current service unit, querying service units other than the current service unit to obtain a target service unit, wherein the target service unit comprises an available memory used for processing the feature blocks; calling the current service unit to process a first feature block among the plurality of feature blocks, so as to obtain a first processing result, and calling the target service unit to process a second feature block among the plurality of feature blocks, so as to obtain a second processing result; and on the basis of the first processing result and the second processing result, generating a task processing result of the target task data. By dividing feature blocks for distributed processing, target task data with a longer context length can be supported, thereby improving the flexibility and adaptability of task processing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This disclosure claims priority to Chinese patent application number 202410010610.3, filed with the Patent Office of the People's Republic of China on January 3, 2024, entitled "Task Processing Method and Automatic Question Answering Method," the entire contents of which are incorporated herein by reference. Technical Field: The embodiments of this disclosure relate to the field of computer technology, and more particularly to task processing methods and automatic question answering methods. Background: With the advancement of computer technology, big models have begun to shine, demonstrating remarkable capabilities in language understanding, generation, interaction, and reasoning, and are widely used in processing fields such as dialogue, translation, and code generation. The rapid development of big models has gradually become a driving force behind the growth of cloud-based big model services, which have become an important component in promoting the application of artificial intelligence. However, due to the dynamic, autoregressive nature of large model services, the lifecycle and sequence length of large model task processing cannot be predetermined. Consequently, the ability of large models to accommodate long context lengths during task processing is limited, resulting in extremely low flexibility and adaptability in the task processing process. Therefore, a highly flexible and adaptable task processing solution is urgently needed. In view of this, embodiments of the present disclosure provide a task processing method. One or more embodiments of the present disclosure also relate to an automatic question-answering method, a task processing apparatus, an automatic question-answering apparatus, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art. According to a first aspect of an embodiment of the present disclosure, a task processing method is provided, comprising: obtaining a feature block sequence of target task data, wherein the feature block sequence includes multiple feature blocks; when a current service unit has insufficient memory, querying a target service unit from a service unit other than the current service unit, wherein the target service unit includes available memory for processing the feature blocks; calling the current service unit to process a first feature block among the multiple feature blocks to obtain a first processing result, and calling the target service unit to process a second feature block among the multiple feature blocks to obtain a second processing result; and generating a task processing result for the target task data based on the first processing result and the second processing result.According to a second aspect of an embodiment of the present disclosure, an automatic question-answering method is provided, comprising: obtaining a feature block sequence for a question to be answered, wherein the feature block sequence includes multiple feature blocks; if a current service unit has insufficient memory, querying a target service unit from a service unit other than the current service unit, wherein the target service unit includes available memory for processing the feature blocks; invoking the current service unit to process a first feature block among the multiple feature blocks to obtain a first processing result, and invoking the target service unit to process a second feature block among the multiple feature blocks to obtain a second processing result; and generating a response to the question to be answered based on the first processing result and the second processing result. According to a third aspect of an embodiment of the present disclosure, a task processing device is provided, comprising: a first acquisition module configured to acquire a feature block sequence of target task data, wherein the feature block sequence includes multiple feature blocks; a first query module configured to, when the current service unit has insufficient memory, query a target service unit from service units other than the current service unit, wherein the target service unit includes available memory for processing feature blocks; a first processing module configured to call the current service unit to process a first feature block among the multiple feature blocks to obtain a first processing result, and call the target service unit to process a second feature block among the multiple feature blocks to obtain a second processing result; and a first generation module configured to generate a task processing result of the target task data based on the first processing result and the second processing result. According to a fourth aspect of an embodiment of the present disclosure, an automatic question-answering apparatus is provided, comprising: a second acquisition module configured to acquire a feature block sequence for a question to be answered, wherein the feature block sequence includes multiple feature blocks; a second query module configured to, if the current service unit has insufficient memory, query a target service unit from a service unit other than the current service unit, wherein the target service unit includes available memory for processing feature blocks; a second processing module configured to invoke the current service unit to process a first feature block among the multiple feature blocks to obtain a first processing result, and invoke the target service unit to process a second feature block among the multiple feature blocks to obtain a second processing result; and a second generation module configured to generate an answer to the question to be answered based on the first processing result and the second processing result. According to a fifth aspect of an embodiment of the present disclosure, a computing device is provided, comprising: a memory and a processor; the memory is configured to store computer-executable instructions, and the processor is configured to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method provided in the first or second aspect.According to a sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, storing computer-executable instructions. When executed by a processor, the instructions implement the steps of the method provided in the first or second aspect. According to a seventh aspect of an embodiment of the present disclosure, a computer program is provided. When executed on a computer, the computer is caused to execute the steps of the method provided in the first or second aspect. A task processing method provided in one embodiment of the present disclosure includes: obtaining a feature block sequence of target task data, wherein the feature block sequence includes multiple feature blocks; when a current service unit has insufficient memory, querying a target service unit from service units other than the current service unit, wherein the target service unit includes available memory for processing feature blocks; invoking the current service unit to process a first feature block among the multiple feature blocks to obtain a first processing result, and invoking the target service unit to process a second feature block among the multiple feature blocks to obtain a second processing result; and generating a task processing result for the target task data based on the first processing result and the second processing result. By invoking the target service unit for processing, underutilized resources in the target service unit are fully utilized, improving resource utilization. Furthermore, by dividing multiple feature blocks into first and second feature blocks for distributed processing, target task data with longer context lengths can be supported, improving the flexibility and adaptability of task processing. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 is an architectural diagram of a task processing system according to one embodiment of the present disclosure; Figure 2 is an architectural diagram of another task processing system according to one embodiment of the present disclosure; Figure 3 is a flow chart of a task processing method according to one embodiment of the present disclosure; Figure 4 is a flow chart of the processing process of a task processing method according to one embodiment of the present disclosure; Figure 5 is a flow chart of an automatic question-and-answer method according to one embodiment of the present disclosure; Figure 6 is a schematic diagram of an automatic question-and-answer interface according to one embodiment of the present disclosure; Figure 7 is a schematic diagram of the structure of a task processing device according to one embodiment of the present disclosure; Figure 8 is a schematic diagram of the structure of an automatic question-and-answer device according to one embodiment of the present disclosure; and Figure 9 is a block diagram of the structure of a computing device according to one embodiment of the present disclosure. The following description sets forth numerous specific details to facilitate a thorough understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art may make similar generalizations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below. The terminology used in one or more embodiments of the present disclosure is for the purpose of describing specific embodiments only and is not intended to limit the one or more embodiments of the present disclosure.As used in one or more embodiments of the present disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should be understood that although the terms first, second, etc. may be employed in one or more embodiments of the present disclosure to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, the first could be referred to as the second, and similarly, the second could be referred to as the first, without departing from the scope of one or more embodiments of the present disclosure. Depending on the context, the word "if," as used herein, could be interpreted as "when..." or "when..." or "in response to determining." Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display) involved in one or more embodiments of the present disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. Large models can also be referred to as foundation models or foundation models. These models are pre-trained using large-scale unlabeled corpora to produce pre-trained models with parameters exceeding 100 million. Such models are adaptable to a wide range of downstream tasks and have good generalization capabilities. Examples include large language models (LLMs) and multimodal pre-training models.In practical applications, large models only require a small number of samples to fine-tune the pre-trained model and can be applied to various tasks. Large models can be widely used in fields such as natural language processing (NLP) and computer vision. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image captioning (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. Key application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design. First, the terms used in one or more embodiments of this disclosure are explained. Distributed Attention Mechanism: The distributed attention mechanism (DAT) is used to improve the efficiency and performance of models when processing large amounts of data. Distributed large-scale language model service system: Distributed Key-Va I ue Large Language Model I (D i stKVLLM) is a large-scale language model service system based on a distributed architecture that can support large-scale user requests and process highly concurrent requests.

[0002] KV Cache: A KV Cache (Key-Va ue Cache) is a distributed caching system used to store and manage data, improving data access speed and efficiency. Cross-data center: Cross-data center refers to the ability of the service system to transmit and process data across multiple data centers, improving system availability and reliability. Graphics Processing Unit (GPU): Graphics Processing Unit (GPU) is used to accelerate computationally intensive tasks such as deep learning. Core Processor Memory: Core Processor Memory refers to the memory within the core processor (CPU) used to store and process data. End-to-end Throughput: End-to-end throughput refers to the processing speed and efficiency of the service system from user request to response result. Context Length: Context length refers to the maximum length of text that the model can process, and is used to improve model accuracy and reliability. Performance Improvement: Performance improvement refers to the performance improvement of the service system when processing large amounts of data and high-concurrency requests, enabling it to better meet user needs. Large-scale language models have driven the rapid growth of LLM services and have become critical infrastructure for the development of AI applications. However, this development faces significant challenges due to its massive computational and data requirements. These services typically utilize multiple GPUs to collaboratively complete LLM tasks. However, the dynamic nature of LLM presents complex computational challenges. At the core of LLM services lies the inherent process of autoregressive text generation, where the model generates one word (or token) at a time. Each newly generated token is appended to the existing text corpus, forming the input for internal LLM calibration. This iterative process continues until the final word or token is generated. Crucially, the memory and computational resources required by LLM services fluctuate continuously throughout the LLM service process, and neither their lifetime nor the sequence length can be known in advance. The dynamic and iterative nature of autoregressive text generation makes it impossible to pre-plan resource allocation, posing a substantial challenge in designing efficient LLM service systems on the cloud. In particular, for long-context tasks, the ever-expanding key-value cache can exceed the GPU memory limits of the compute instance, necessitating immediate resource reallocation. This typically involves initiating costly live migrations to move the task to more capable instances, or pre-allocating additional GPUs to handle potential memory overloads. However, the latter approach can lead to inefficiencies and resource waste, especially for tasks with moderate context lengths. Currently, the above problems can usually be solved by facilitating data exchange between GPU and CPU memory. However, this approach encounters several limitations.First, memory swapping is limited to the GPU and CPU memory within a single node, limiting its ability to accommodate extremely long context lengths. Second, because this solution swaps the entire KV cache at the request level, it misses the opportunity for more adaptive, fine-grained scheduling in distributed cloud environments. Finally, computational interruptions caused by swapped requests can cause performance fluctuations in running tasks, potentially failing to meet the stringent service-level agreements (SLAs) critical to cloud services. To address the aforementioned issues, embodiments of the present disclosure propose a task processing method that obtains a feature block sequence for target task data, where the feature block sequence includes multiple feature blocks. When the current service unit is low on memory, a target service unit is retrieved from service units other than the current service unit, where the target service unit includes available memory for processing feature blocks. The current service unit is invoked to process a first feature block from the multiple feature blocks to obtain a first processing result, and the target service unit is invoked to process a second feature block from the multiple feature blocks to obtain a second processing result. Based on the first and second processing results, a task processing result for the target task data is generated. By invoking the target service unit for processing, underutilized resources in the target service unit are fully utilized, improving resource utilization. Furthermore, by dividing multiple feature blocks into first and second feature blocks for distributed processing, target task data with longer context lengths can be supported, improving the flexibility and adaptability of task processing. This disclosure provides a task processing method, an automatic question-answering method, a task processing apparatus, an automatic question-answering apparatus, a computing device, and a computer-readable storage medium, each of which is described in detail in the following embodiments.1 , which shows an architecture diagram of a task processing system provided by an embodiment of the present disclosure. The task processing system may include a client 100 and a server 200, wherein the server 200 includes multiple service units. The client 100 is configured to send a feature block sequence of target task data to the server 200, wherein the feature block sequence includes multiple feature blocks. The server 200 is configured to query a target service unit from a service unit other than the current service unit when the current service unit has insufficient memory, wherein the target service unit includes available memory for processing feature blocks. The current service unit is invoked to process a first feature block among the multiple feature blocks to obtain a first processing result, and the target service unit is invoked to process a second feature block among the multiple feature blocks to obtain a second processing result. A task processing result for the target task data is generated based on the first processing result and the second processing result. The task processing result is sent to the client 100. The client 100 is also configured to receive the task processing result sent by the server 200. Applying the solution of the embodiments of the present disclosure, by invoking the target service unit for processing, fully utilizes underutilized resources in the target service unit, improving resource utilization. Furthermore, by dividing multiple feature blocks into first and second feature blocks for distributed processing, it can support target task data with longer context lengths, thereby improving the flexibility and adaptability of task processing. Referring to Figure 2, Figure 2 shows an architecture diagram of another task processing system provided by one embodiment of the present disclosure. The task processing system may include multiple clients 100 and a server 200. The clients 100 may include end-side devices, and the server 200 may include cloud-side devices. Multiple clients 100 can establish communication connections through the server 200. In a task processing scenario, the server 200 is used to provide task processing services between the multiple clients 100. The multiple clients 100 can act as senders or receivers, respectively, and communicate through the server 200. Users can interact with the server 200 through the clients 100 to receive data from other clients 100 or send data to other clients 100. In a task processing scenario, a user may publish a data stream to a server 200 via a client 100. The server 200 generates a task processing result based on the data stream and pushes the result to other clients with which the communication is established. The connection between the client 100 and the server 200 is established via a network. The network provides the medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables.The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or otherwise processed before being published to the server 200. The client 100 may be a browser, an App (Application Program), a web application such as an H5 (HyperText Markup Languages, version 5) application, a lightweight application (also known as a mini-program, a type of lightweight application), or a cloud application. The client 100 may be developed based on a software development kit (SDK) for a corresponding service provided by the server 200, such as a real-time communication (RTC) SDK. The client 100 may be deployed in an electronic device and may rely on the device or certain apps in the device to operate. For example, the electronic device may have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, or a personal computer. Electronic devices can also typically be configured with various other types of applications, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, and the like. The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that provide backend training support for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers or as a single server. The server can also be a server in a distributed system or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data and artificial intelligence platforms, or intelligent cloud computing servers or intelligent cloud hosts with artificial intelligence technology. It is worth noting that the task processing methods provided in the embodiments of the present disclosure are generally executed by the server. However, in other embodiments of the present disclosure, the client may also have similar functions to the server and thus execute the task processing methods provided in the embodiments of the present disclosure. In other embodiments, the task processing methods provided in the embodiments of the present disclosure may also be jointly executed by the client and the server.Referring to Figure 3, a flowchart of a task processing method provided by one embodiment of the present disclosure is shown. The method specifically includes the following steps: Step 302: Obtain a feature block sequence of target task data, where the feature block sequence includes multiple feature blocks. In one or more embodiments of the present disclosure, to improve the flexibility and adaptability of task processing, a feature block sequence of target task data may be obtained, and distributed task processing may be performed based on the feature block sequence. Specifically, target task data refers to the data to be processed corresponding to the target task. Target tasks may be different tasks in different scenarios, such as e-commerce and conference scenarios, or different tasks such as question-and-answer tasks, retrieval tasks, and summary extraction tasks. The data to be processed may be data of different modalities, such as voice data to be processed, text data to be processed, or image data to be processed. It should be noted that to facilitate distributed task processing, the KV cache may be divided into smaller units, referred to as sub-blocks (rBlocks). A feature block refers to a sub-block that stores the task data features of the target task data. In practical applications, there are various methods for obtaining a feature block sequence for target task data, and the selection should be based on practical circumstances. The present disclosure does not impose any restrictions on this method. In one possible implementation of the present disclosure, the feature block sequence for target task data can be read from another data acquisition device or database. In another possible implementation of the present disclosure, obtaining the feature block sequence for target task data may include the following steps: obtaining target task data; performing feature extraction on the target task data to obtain a task feature sequence, where the task feature sequence includes multiple task data features; and constructing a feature block sequence based on a feature block partitioning condition and the multiple task data features. Specifically, the feature block partitioning condition is used to partition the multiple task data features into at least one feature block. The feature block partitioning condition includes, but is not limited to, the number of task data features in a feature block and the number of feature blocks in a feature block sequence, and the selection should be based on practical circumstances. The present disclosure does not impose any restrictions on this method. It should be noted that there are various methods for obtaining target task data, and the selection should be based on practical circumstances. The present disclosure does not impose any restrictions on this method. In one possible implementation of the present disclosure, target task data sent by a user can be received. In another possible implementation of the present disclosure, target task data can be read from other data acquisition devices or databases. Furthermore, there are multiple ways to extract features from the target task data to obtain a task feature sequence, and the method can be selected based on actual circumstances. This disclosure does not impose any limitations on this method.In one possible implementation of the present disclosure, the target task data can be input into a deep learning model (such as a convolutional neural network or a recurrent neural network) for feature extraction to obtain a task feature sequence. In another possible implementation of the present disclosure, a word embedding model (such as Word2Vec) can be used to extract features from the target task data to obtain a task feature sequence. For example, assume that the task feature sequence includes 12 task data features, and the feature block division condition is that the number of task data features in a feature block is 4. Then, the 1st to 4th task data features can be stored in sub-block 1 to obtain feature block 1, the 5th to 8th task data features can be stored in sub-block 2 to obtain feature block 2, and the 9th to 12th task data features can be stored in sub-block 3 to obtain feature block 3. A feature block sequence can be constructed based on feature blocks 1, 2, and 3. Applying the solutions of the embodiments of the present disclosure, target task data is acquired; features are extracted from the target task data to obtain a task feature sequence, where the task feature sequence includes multiple task data features; and a feature block sequence is constructed based on the feature block partitioning conditions and the multiple task data features. This effectively decomposes the KV cache into smaller, more manageable sub-blocks, thereby facilitating distributed memory management across data centers. Step 304: If the current service unit is short of memory, a target service unit is retrieved from service units other than the current service unit. The target service unit includes available memory for processing the feature blocks. In one or more embodiments of the present disclosure, after acquiring the feature block sequence of the target task data, it is further determined whether the current service unit has sufficient memory. If the current service unit is short of memory, the target service unit is retrieved from service units other than the current service unit. Specifically, the current service unit is a service unit among multiple service units in the task processing system that is used to process the target task data. Therefore, the current service unit can be referred to as a local service unit or local instance of the task processing. Service units other than the current service unit can be referred to as remote service units or remote instances of the task processing. The target service unit includes available memory for processing feature blocks. Therefore, the target service unit can be understood as an example of excess capacity. It should be noted that if the current service unit is short on memory, it indicates that the current service unit cannot complete processing of multiple feature blocks. In this case, available memory space can be borrowed from service units with excess capacity to collaboratively complete processing of multiple feature blocks. In other words, the target service unit is obtained by querying service units other than the current service unit.In an optional embodiment of the present disclosure, before querying for a target service unit from a service unit other than the current service unit, it is possible to determine whether the current service unit's memory is insufficient. If the current service unit has sufficient memory to process multiple feature blocks, the current service unit can be directly invoked to process the multiple feature blocks, obtain processing results, and generate a task processing result for the target task data based on the processing results. If the current service unit does not have sufficient memory to process the multiple feature blocks, the target service unit can be queried from a service unit other than the current service unit. In practical applications, there are various ways to determine whether the current service unit's memory is insufficient, and the specific method to be used depends on the actual situation. This embodiment of the present disclosure does not impose any limitation on this method. In one possible implementation of the present disclosure, the current service unit can be invoked to process multiple feature blocks, and the memory usage of the current service unit can be monitored in real time. If feature block processing fails, it is determined that the current service unit has insufficient memory. In another possible implementation of the present disclosure, whether the current service unit's memory is insufficient can be determined based on the processing memory corresponding to multiple feature blocks. Specifically, before querying for a target service unit from service units other than the current service unit in the case of insufficient memory, the following steps may be included: obtaining the current available memory of the current service unit and determining the processing memory corresponding to the multiple feature blocks; and determining that the current service unit's memory is insufficient if the current available memory is less than the processing memory. Specifically, the current available memory refers to unoccupied memory space that can process feature blocks. The current available memory can also be understood as the free memory of the current service unit. The processing memory refers to the memory space required to process feature blocks. In actual applications, there are various methods for obtaining the current available memory of the current service unit and determining the processing memory corresponding to the multiple feature blocks, and the methods can be selected based on actual circumstances. This embodiment of the present disclosure does not impose any limitation on this. In one possible implementation of the present disclosure, current available memory and processing memory information sent by a user can be received. In another possible implementation of the present disclosure, the current available memory of the current service unit and the processing memory corresponding to the multiple feature blocks can be read from a task manager. It should be noted that after obtaining the current available memory of the current service unit and determining the processing memory corresponding to multiple feature blocks, the memory sizes of the current available memory and the processing memory can be compared. If the current available memory is greater than or equal to the processing memory, it means that the current service unit has sufficient memory and can independently complete task processing. If the current available memory is less than the processing memory, it means that the current service unit has insufficient memory and cannot independently process the task. It needs to borrow available memory space from service units with excess capacity and collaborate with other service units to complete the processing of multiple feature blocks.Using the solution of the embodiments of the present disclosure, the currently available memory of the current service unit is obtained, and the processing memory corresponding to multiple feature blocks is determined. If the currently available memory is less than the processing memory, the current service unit is determined to have insufficient memory. By determining whether the current service unit has insufficient memory based on the current available memory and the processing memory, interruptions in task processing due to insufficient service unit memory are avoided, ensuring the continuity of task processing. In actual applications, there are various methods for querying and obtaining a target service unit from service units other than the current service unit. The specific method to be used depends on the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the processing memory of multiple feature blocks can be evenly divided. The target service unit can be selected from the multiple service units based on the divided processing memory. The current service unit and the target service unit can then process the multiple feature blocks equally. In another possible implementation of the present disclosure, a target service unit can be queried from service units other than the current service unit based on the memory difference between the current available memory and the processing memory corresponding to multiple feature blocks, thereby fully utilizing the available memory of the current service unit. Specifically, the aforementioned process of querying for a target service unit from service units other than the current service unit may include the following steps: determining memory requirement information based on the memory difference between the current available memory of the current service unit and the processing memory corresponding to multiple feature blocks; screening multiple candidate service units from the multiple service units based on the unit memory usage table and the memory requirement information, wherein the candidate service units include available memory that meets the memory requirement information; and screening the target service unit from the multiple candidate service units. It should be noted that the memory requirement information indicates the amount of memory space required by the current service unit. For example, if the current available memory is a and the processing memory is b, the memory requirement information is ba. The unit memory usage table is used to track and record the memory usage of each service unit. The unit memory usage table includes the available memory information of the service unit and memory interaction information between the service units. The memory interaction information includes the service unit ID, the memory interaction size, and the memory interaction relationship. For example, service unit 1 has 80 GB of memory and 30 GB of available memory. Service unit 1 borrows 11 GB of memory from service unit 2 and borrows 20 GB of memory from service unit 3.For example, assuming that the unit memory usage table includes service unit 1, service unit 2, and service unit 3, the available memory of service unit 1 is 30 GB, the available memory of service unit 2 is 3 GB, the available memory of service unit 3 is 50 GB, and the memory requirement information is 18 GB, then based on the unit memory usage table and the memory requirement information, multiple candidate service units are screened from multiple service units to be service unit 1 and service unit 3. OIn practical applications, the task processing system can be understood as a distributed large-scale language model service system. Within this distributed large-scale language model service system, a management unit (gManager) can be set up, and each service unit can be equipped with a processing unit (rManager). Multiple service units share one management unit. The processing unit virtualizes the GPU and CPU memory within the service unit and provides a unified application programming interface (API) to service local and remote memory operations, including allocating memory for feature blocks and releasing memory when no longer needed for processing. The management unit acts as a global coordinator, maintaining global memory information for all service units and ensuring efficient, scalable, and consistent resource management across the distributed service units. When a service unit is running low on memory, the rManager of the current service unit attempts to borrow GPU or CPU memory space from neighboring service units. This involves initiating a query to the gManager and sending it memory request information. After receiving the query, gManager can query the unit memory usage table and provide the address ID of a candidate service unit. This ID represents a service unit with available memory. rManager then selects a target service unit from multiple candidate service units. Using the solution of the embodiments of the present disclosure, memory requirements are determined based on the difference between the current available memory of the current service unit and the processing memory corresponding to multiple feature blocks. Based on the unit memory usage table and the memory requirements, multiple candidate service units are selected from multiple service units, where the candidate service units include available memory that meets the memory requirements. The target service unit is selected from the multiple candidate service units, thereby avoiding selecting a service unit with insufficient available memory as the target service unit, preventing service unit overload, and ensuring the availability of the target service unit. In practical applications, there are various methods for selecting a target service unit from multiple candidate service units, and the specific method is selected based on actual circumstances. This embodiment of the present disclosure does not impose any limitations on this method. In one possible implementation of the present disclosure, the target service unit can be randomly selected from the multiple candidate service units.In another possible implementation of the present disclosure, a target service unit can be selected from multiple candidate service units based on the principles of locality and availability. Specifically, selecting the target service unit from the multiple candidate service units may include the following steps: sending a memory borrowing request to the candidate service units based on their priorities; upon receiving borrowing confirmation information returned by a first candidate service unit, determining the first candidate service unit as the target service unit, where the first candidate service unit is any one of the multiple candidate service units. It should be noted that the priorities of the multiple candidate service units may be determined based on communication costs and / or available memory size. The multiple candidate service units may be sorted based on their priorities to determine a request sending queue. Furthermore, based on the request sending queue, a memory borrowing request may be sent to a higher-priority candidate service unit. If the higher-priority candidate service unit returns a borrowing confirmation information, the candidate service unit may be determined as the target service unit. If the higher-priority candidate service unit returns a borrowing rejection information, the memory borrowing request may be sent to the next lower-priority candidate service unit until the target service unit is found. In practical applications, if a service unit receives memory borrowing requests from multiple remote service units simultaneously, it can borrow memory for the remote service units based on a first-come, first-served policy. If a candidate service unit returns a borrowing rejection, forwarding memory borrowing requests to that candidate service unit can be suspended until more memory becomes available. Using the solution of the disclosed embodiments, memory borrowing requests are sent to candidate service units based on their priorities. Upon receiving a borrowing confirmation from the first candidate service unit, the first candidate service unit is identified as the target service unit. This ensures balanced and orderly allocation of memory resources, selects the target service unit from multiple candidate service units based on lower relative communication costs and higher available memory space, and alleviates potential task processing bottlenecks in the system. In an optional embodiment of the present disclosure, before selecting multiple candidate service units from multiple service units based on the unit memory usage table and memory requirement information, the following steps may be further included: obtaining heartbeat information of the multiple service units, where the heartbeat information includes available memory information and memory interaction information of the service units; and constructing a unit memory usage table based on the available memory information and memory interaction information of the multiple service units. Specifically, the available memory information includes, but is not limited to, available memory addresses and available memory sizes. The memory interaction information includes, but is not limited to, memory interaction sizes and memory interaction relationships. Memory interaction includes, but is not limited to, borrowing and lending memory between service units.In practical applications, the rManager can periodically send service unit heartbeat information to the gManager, which then maintains a unit memory usage table summarizing the service unit's memory usage. This eliminates the need for the gManager to meticulously track every memory allocation or release operation across all service units. Using the solution of the disclosed embodiments, heartbeat information from multiple service units is obtained, including each service unit's available memory and memory interaction information. Based on this information, a unit memory usage table is constructed, streamlining the memory space borrowing process, reducing processing latency, and improving overall system performance. In practical applications, each service unit in the system can act as a memory creditor and debtor, interacting with other service units as needed. For example, a service unit processing requests with long contexts may grow in size and therefore need to borrow space from remote service units. Conversely, service units with short-lived requests release memory space more quickly, which can then be loaned to other service units or allocated for new requests. This dynamic nature can lead to deterioration of data locality: because service units frequently access data stored in remote memory locations, the system can suffer significant performance losses, such as increased latency and reduced end-to-end throughput. Therefore, in the embodiments of this disclosure, a fragmented memory management strategy is proposed. This strategy aims to offset memory interactions by strategically recalling loaned memory space and exchanging local feature blocks based on memory interaction relationships. Specifically, after obtaining heartbeat information from multiple service units, the following steps may also be included: parsing the memory interaction information to determine the memory interaction relationships between the multiple service units; constructing a memory interaction graph with the service units as nodes and the memory interaction relationships as edges; and adjusting the feature blocks corresponding to the service units based on the memory interaction graph. Specifically, memory interaction relationships refer to memory borrowing or lending relationships between multiple service units. The memory interaction graph is a directed graph, meaning that an edge in the memory interaction graph represents a directed edge representing memory interaction from one service unit to another. In practical applications, there are various ways to parse memory interaction information and determine the memory interaction relationships between multiple service units. The method to be used depends on the specific situation and is not limited in this embodiment of this disclosure. In one possible implementation of the present disclosure, the memory interaction relationship can be parsed from the memory logs of multiple service units. In another possible implementation of the present disclosure, the memory interaction relationship can be parsed from the memory interaction information using a resource management and monitoring tool.It should be noted that there are multiple ways to adjust the feature blocks corresponding to service units based on the memory interaction graph, and the specific method to be used depends on actual circumstances. This disclosure does not impose any restrictions on this method. In one possible implementation of this disclosure, the feature blocks corresponding to service units can be adjusted directly based on the memory interaction graph. In another possible implementation of this disclosure, to avoid adjustment failure caused by continued memory interaction between the service units during the adjustment process, the service units can be frozen first, and then the feature blocks corresponding to the service units can be adjusted based on the memory interaction graph. Using the solution of this embodiment of the disclosure, memory interaction information is parsed to determine the memory interaction relationships between multiple service units; a memory interaction graph is constructed with the service units as nodes and the memory interaction relationships as edges; and the feature blocks corresponding to the service units are adjusted based on the memory interaction graph. By adjusting the feature blocks corresponding to each service unit, overall system efficiency and memory utilization are improved. In an optional embodiment of the present disclosure, adjusting the feature blocks corresponding to service units based on the memory interaction graph may include the following steps: selecting the service units to be adjusted from the memory interaction graph, wherein the service units to be adjusted form a ring relationship structure in the memory interaction graph; and adjusting the feature blocks corresponding to the service units to be adjusted based on the memory interaction information of the service units to be adjusted. It should be noted that when selecting the service units to be adjusted from the memory interaction graph, a node can be randomly selected in the memory interaction graph, and the memory interaction graph can be traversed to find at least one ring relationship structure. This ring relationship structure indicates that the nodes involved can offset each other's interaction memory. For example, assume that service unit 1 borrows feature block 1 from service unit 2, service unit 2 borrows feature block 2 from service unit 3, and service unit 3 borrows feature block 3 from service unit 1. In this case, a ring relationship structure is formed between these three service units. In this case, feature block 1 can be returned and stored in service unit 1, feature block 2 can be returned and stored in service unit 2, and feature block 3 can be returned and stored in service unit 3, thereby offsetting the interaction memory between them. Using the solution of the embodiments of the present disclosure, service units to be adjusted are screened from a memory interaction graph, where the service units to be adjusted form a ring relationship structure in the memory interaction graph. Based on the memory interaction information of the service units to be adjusted, the corresponding feature blocks of the service units to be adjusted are adjusted. This reduces the need for memory access to remote service units, thereby improving data locality and ultimately significantly improving system performance. Step 306: Invoke the current service unit to process a first feature block from the multiple feature blocks to obtain a first processing result, and then invoke the target service unit to process a second feature block from the multiple feature blocks to obtain a second processing result.In one or more embodiments of the present disclosure, a sequence of feature blocks of target task data is obtained. If the current service unit has insufficient memory, a target service unit is obtained by querying a service unit other than the current service unit. Furthermore, the current service unit may be called to process a first feature block among multiple feature blocks to obtain a first processing result, and the target service unit may be called to process a second feature block among the multiple feature blocks to obtain a second processing result. Specifically, when calling a service unit to process a feature block, a processing result may be generated based on a distributed attention mechanism. That is, when calling the current service unit to process a first feature block among multiple feature blocks, the current service unit may be called to perform attention calculation on the first feature block among the multiple feature blocks to obtain a first attention result. When calling the target service unit to process a second feature block among the multiple feature blocks, the target service unit may be called to perform attention calculation on the second feature block among the multiple feature blocks to obtain a second attention result. The attention mechanism is used to dynamically assign attention weights when the model processes input data. The attention calculation process typically includes the calculation of key vectors, value vectors, and query vectors; the calculation of attention weights, where the weights reflect the correlation between the query vector and each key vector, that is, the importance of each part of the input sequence to the current task; the aggregation of attention weights: applying the attention weights to the corresponding value vectors; and the generation of attention results. In an optional embodiment of the present disclosure, due to insufficient memory in the current service unit, the current service unit and the target service unit can collaboratively process multiple feature blocks. In this case, the original processing process can be divided into two micro-processes based on the number of service units participating in the attention processing, thereby completing the processing in a distributed manner. During each microprocessing, its corresponding feature sequence needs to be determined. Specifically, before calling the current service unit to process a first feature block among the multiple feature blocks to obtain a first processing result and calling the target service unit to process a second feature block among the multiple feature blocks to obtain the second processing result, the following steps may also be included: obtaining currently available memory of the current service unit; and dividing the multiple feature blocks according to the currently available memory to determine a first feature block and a second feature block, wherein the currently available memory is used to process the first feature block. For example, assume there are 2000 feature blocks. The currently available memory of the current service unit is 30 GB, which can be used to process 1000 feature blocks.In this case, the multiple feature blocks can be divided based on the currently available memory, with the first 1000 feature blocks determined as the first feature block and the last 1000 feature blocks determined as the second feature block. Applying the solution of the embodiment of the present disclosure, the currently available memory of the current service unit is obtained; the multiple feature blocks are divided based on the currently available memory to determine the first feature block and the second feature block, thereby ensuring accurate distributed processing. In an optional embodiment of the present disclosure, the aforementioned invoking of the target service unit to process the second feature block among the multiple feature blocks and obtain the second processing result may include the following steps: searching the target memory address from the address mapping table of the target service unit based on the attribute information of the second feature block; storing the second feature block in the target service unit based on the target memory address, and invoking the target service unit to process the second feature block. Specifically, the attribute information of the second feature block can be understood as metadata of the second feature block, which is used to provide logical sub-block information of the relevant sub-block. The address mapping table includes a mapping relationship between logical sub-block information and physical sub-block information. The memory address can be understood as a physical address. It should be noted that the implementation method for invoking the current service unit to process the first feature block among multiple feature blocks is the same as the implementation method for invoking the target service unit to process the second feature block among multiple feature blocks to obtain the second processing result, and will not be further described in detail in this embodiment of the present disclosure. Referring to FIG4 , FIG4 shows a flowchart of a task processing method provided by one embodiment of the present disclosure, specifically including the following: During task processing, the processing units in the local service unit and the remote service unit are used to divide the CPU and multiple GPUs into physical sub-blocks of fixed size, effectively virtualizing global memory. The processing units in the multiple service units share a management unit. Furthermore, the processing units can allocate memory for feature blocks and release memory when feature blocks are no longer needed for processing. The processing units are responsible for maintaining an address mapping table to map logical sub-blocks to corresponding physical sub-blocks in global memory. As shown in FIG4 , the logical sub-block information includes a sub-block ID (Identity Document) and a service unit ID. The sub-block ID and service unit ID indicate whether the KV cache for the feature block is stored in the local service unit or the remote service unit. The physical sub-block information includes a device identifier and a physical address identifier. The device identifier and physical address identifier are used to indicate whether the physical memory address of the feature block is located on the CPU side or on one of the multiple GPUs. In practical applications, for any service unit, the feature block can be processed using the following formula (1):

[0003] MA" = exp(QjKj - max(Q占; )) (1) Where, MA" represents the processing result; exp represents the exponential function; Q represents the query vector; K represents the key vector; T represents matrix transpose; i represents the i-th token in the feature block corresponding to the service unit; j represents the j-th key vector segmentation block of the feature block; QKj represents the inner product result obtained by taking the inner product of the query vector of the i-th token and the j-th key vector segmentation block; max(Q〔K:) represents the maximum value of the inner product result. Applying the solution of the embodiment of the present disclosure, according to the attribute information of the second feature block, the target memory address is searched from the address mapping table of the target service unit; according to the target memory address, the second feature block is stored in the target service unit, and the target service unit is called to process the second feature block, realizing distributed processing, so as to support target task data with a longer context length, and improving the flexibility and adaptability of task processing. Step 308: Generate the task processing result of the target task data according to the first processing result and the second processing result. In one or more embodiments of the present disclosure, a sequence of feature blocks of the target task data is obtained; when the memory of the current service unit is insufficient, the target service unit is queried from the service units other than the current service unit; the current service unit is called to process the first feature block in multiple feature blocks to obtain the first processing result, and the target service unit is called to process the second feature block in multiple feature blocks to obtain the second processing result. Further, the task processing result of the target task data can be generated according to the first processing result and the second processing result. Specifically, the task processing result is related to the task corresponding to the target task data. For example, if the task is a question-and-answer task, the task processing result is the answer result; if the task is a retrieval task, the task processing result is the retrieval result. Applying the solution of the embodiment of the present disclosure, by calling the target service unit for processing, the resources not fully utilized in the target service unit are fully utilized, improving the resource utilization rate. And by dividing multiple feature blocks into the first feature block and the second feature block for distributed processing, it can support target task data with a longer context length, improving the flexibility and adaptability of task processing. At the same time, it avoids performance fluctuations related to data exchange or real-time migration processes. In practical applications, there are various ways to generate the task processing result of the target task data according to the first processing result and the second processing result, which are specifically selected according to the actual situation, and the embodiments of the present disclosure do not make any limitations on this.In one possible implementation of the present disclosure, a first task processing result can be generated based on the first processing result, a second task processing result can be generated based on the second processing result, and a task processing result for target task data can be generated based on the first task processing result and the second task processing result. In another possible implementation of the present disclosure, the above-mentioned generation of the task processing result for target task data based on the first processing result and the second processing result can include the following steps: aggregating the first processing result and the second processing result to obtain an aggregated processing result; and generating the task processing result for target task data based on the aggregated processing result. In practical applications, aggregation refers to performing aggregate processing on distributed processing results. During aggregation, the first processing result and the second processing result can be scaled and added together to obtain an aggregated processing result. Specifically, the first processing result and the second processing result can be aggregated using the following formulas (2), (3), and (4) to obtain an aggregated processing result:

[0004] Atention(Q,K,V) = Reduce(Scale([MA; j ]Jf ))

[0005] = Reduce([ex p(max(Q ; Kj) - max ; )MA Gongfuj ) (2)

[0006] = Z (exp(max( Q ; Kj) - max ; )MA; j / sum ;) j=i maXi = max(max(Q jK: max(QiK=)) (3) sunii = Z (exp(QjK:) — maXi) (4) Wherein, Atention( Q,K,V) represents the aggregation processing result; Reduce represents the aggregation operation; Scale represents the scaling operation; Bkv represents the number of blocks into which the key vector and the value vector are cut; maXi represents the maximum value of the attention weight corresponding to the i-th query vector; QX: represents the inner product of the query vector of the i-th token and the 1st key vector cut block; QiK= represents the inner product of the query vector of the i-th token and the B-th key-value vector cut block; sunii represents the denominator of the softmax calculation corresponding to the query vector of the i-th token. Further, after obtaining the aggregation processing result, the context vector can be obtained by weighted summation according to the aggregation processing result, and the context vector can be decoded to obtain the task processing result of the target task data. Applying the solution of the embodiment of the present disclosure, the first processing result and the second processing result are aggregated to obtain the aggregation processing result; according to the aggregation processing result, Generate task processing results for target task data. Through distributed processing, target task data with longer context lengths can be supported, thereby improving the flexibility and adaptability of task processing. The following, in conjunction with Figure 5, takes the application of the task processing method provided by the present disclosure in an intelligent question-answering scenario as an example to further illustrate the task processing method. Figure 5 shows a flowchart of an automatic question-answering method provided by an embodiment of the present disclosure, which specifically includes the following steps: Step 502: Obtain a feature block sequence of the question to be answered, wherein the feature block sequence includes multiple feature blocks. Step 504: When the current service unit has insufficient memory, query a target service unit from a service unit other than the current service unit, wherein the target service unit includes available memory for processing feature blocks. Step 506: Call the current service unit to process a first feature block among the multiple feature blocks to obtain a first processing result, and call the target service unit to process a second feature block among the multiple feature blocks to obtain a second processing result. Step 508: Generate a response result for the question to be answered based on the first processing result and the second processing result. It should be noted that The implementation of steps 502 to 508 may refer to the implementation of steps 302 to 308 described above, and will not be described in detail in this embodiment of the disclosure.Applying the solution of the embodiments of the present disclosure, by invoking the target service unit for processing, fully utilizes underutilized resources in the target service unit, improving resource utilization. Furthermore, by dividing multiple feature blocks into first and second feature blocks for distributed processing, it can support questions with longer context lengths, improving the flexibility and adaptability of automatic question answering. See Figure 6, which shows a schematic diagram of an automatic question answering interface provided by one embodiment of the present disclosure. The task processing interface is divided into a request input interface and a result display interface. The request input interface includes a request input box, an "OK" control, and a "Cancel" control. The result display interface includes a result display box. The user enters an automatic question-and-answer request containing a feature block sequence of the question to be answered through the request input box displayed on the client. The feature block sequence includes multiple feature blocks. The user clicks the "OK" control. The server receives the feature block sequence sent by the client. If the current service unit is short of memory, it queries a target service unit from service units other than the current one. The target service unit includes available memory for processing feature blocks. The server then calls the current service unit to process a first feature block from the multiple feature blocks, obtaining a first processing result. The server then calls the target service unit to process a second feature block from the multiple feature blocks, obtaining a second processing result. Based on the first and second processing results, the server generates an answer to the question to be answered and sends the answer to the client. The client displays the answer in the result display box. In actual applications, user operations on the control may include clicking, double-clicking, touching, hovering the mouse, sliding, long pressing, voice control, or shaking. The specific operation is selected based on actual circumstances and is not limited in this embodiment of the disclosure. Corresponding to the above-mentioned task processing method embodiment, the present disclosure further provides a task processing device embodiment. FIG7 shows a schematic structural diagram of a task processing device provided by an embodiment of the present disclosure.As shown in Figure 7, the device includes: a first acquisition module 702, configured to obtain a feature block sequence of target task data, wherein the feature block sequence includes multiple feature blocks; a first query module 704, configured to query a target service unit from a service unit other than the current service unit when the current service unit has insufficient memory, wherein the target service unit includes available memory for processing feature blocks; a first processing module 706, configured to call the current service unit to process a first feature block among the multiple feature blocks to obtain a first processing result, and call the target service unit to process a second feature block among the multiple feature blocks to obtain a second processing result; and a first generation module 708, configured to generate a task processing result of the target task data based on the first processing result and the second processing result. Optionally, the first query module 704 is further configured to determine memory requirement information based on the memory difference between the current available memory of the current service unit and the processing memory corresponding to the multiple feature blocks; filter multiple candidate service units from the multiple service units based on the unit memory usage table and the memory requirement information, wherein the candidate service units include available memory that meets the memory requirement information; and filter a target service unit from the multiple candidate service units. Optionally, the first query module 704 is further configured to send a memory borrowing request to the multiple candidate service units based on their priorities; upon receiving borrowing confirmation information returned by the first candidate service unit, determine the first candidate service unit as the target service unit, wherein the first candidate service unit is any one of the multiple candidate service units. Optionally, the apparatus further includes: a construction module configured to obtain heartbeat information of the multiple service units, wherein the heartbeat information includes the service units' available memory information and memory interaction information; and construct a unit memory usage table based on the available memory information and memory interaction information of the multiple service units. Optionally, the apparatus further includes: an adjustment module configured to parse memory interaction information to determine memory interaction relationships between multiple service units; construct a memory interaction graph using the service units as nodes and the memory interaction relationships as edges; and adjust feature blocks corresponding to the service units based on the memory interaction graph. Optionally, the adjustment module is further configured to filter out service units to be adjusted from the memory interaction graph, wherein the service units to be adjusted form a ring relationship structure in the memory interaction graph; and adjust feature blocks corresponding to the service units to be adjusted based on the memory interaction information of the service units to be adjusted.Optionally, the first processing module 706 is further configured to search the target memory address from the address mapping table of the target service unit based on the attribute information of the second feature block; store the second feature block in the target service unit based on the target memory address, and invoke the target service unit to process the second feature block. Optionally, the first generation module 708 is further configured to aggregate the first processing result and the second processing result to obtain an aggregated processing result; and generate a task processing result for the target task data based on the aggregated processing result. Optionally, the first acquisition module 702 is further configured to acquire target task data; perform feature extraction on the target task data to obtain a task feature sequence, wherein the task feature sequence includes multiple task data features; and construct a feature block sequence based on the feature block partitioning condition and the multiple task data features. Optionally, the apparatus further includes: a partitioning module configured to acquire currently available memory of the current service unit; partition the multiple feature blocks based on the currently available memory to determine a first feature block and a second feature block, wherein the currently available memory is used to process the first feature block. Applying the solution of the embodiments of the present disclosure, by invoking the target service unit for processing, fully utilizes underutilized resources in the target service unit, improving resource utilization. Furthermore, by dividing multiple feature blocks into first and second feature blocks for distributed processing, it is possible to support target task data with longer context lengths, thereby improving the flexibility and adaptability of task processing. The above is a schematic diagram of a task processing device according to this embodiment. It should be noted that the technical solution of this task processing device and the technical solution of the aforementioned task processing method are based on the same concept. For details not described in detail in the technical solution of the task processing device, please refer to the description of the technical solution of the aforementioned task processing method. Corresponding to the aforementioned automatic question-answering method embodiment, the present disclosure also provides an embodiment of an automatic question-answering device. Figure 8 shows a schematic structural diagram of an automatic question-answering device according to one embodiment of the present disclosure.As shown in Figure 8 , the apparatus includes: a second acquisition module 802 configured to acquire a feature block sequence for a question to be answered, wherein the feature block sequence includes multiple feature blocks; a second query module 804 configured to, when the current service unit is short of memory, query a target service unit from a service unit other than the current service unit, wherein the target service unit includes available memory for processing feature blocks; a second processing module 806 configured to invoke the current service unit to process a first feature block from the multiple feature blocks to obtain a first processing result, and invoke the target service unit to process a second feature block from the multiple feature blocks to obtain a second processing result; and a second generation module 808 configured to generate a response to the question to be answered based on the first processing result and the second processing result. By applying the solution of the embodiments of the present disclosure, by invoking the target service unit for processing, underutilized resources in the target service unit are fully utilized, thereby improving resource utilization. Furthermore, by dividing the multiple feature blocks into first and second feature blocks for distributed processing, the system can support questions with longer context lengths, thereby improving the flexibility and adaptability of automatic question answering. The above is a schematic diagram of an automatic question-answering device according to this embodiment. It should be noted that the technical solution of this automatic question-answering device and the technical solution of the automatic question-answering method described above share the same concept. For details not described in detail in the technical solution of the automatic question-answering device, please refer to the description of the technical solution of the automatic question-answering method described above. Figure 9 shows a block diagram of a computing device according to one embodiment of the present disclosure. The components of computing device 900 include, but are not limited to, a memory 910 and a processor 920. Processor 920 is connected to memory 910 via a bus 930, and a database 950 is used to store data. Computing device 900 also includes an access device 940, which enables computing device 900 to communicate via one or more networks 960. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet.The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like. In one embodiment of the present disclosure, the aforementioned components of the computing device 900 and other components not shown in FIG. 9 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG. 9 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed. The computing device 900 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 900 can also be a mobile or stationary server. The processor 920 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the aforementioned task processing method or automatic question-answering method. The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the aforementioned task processing method and automatic question-answering method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned task processing method or automatic question-answering method. An embodiment of the present disclosure further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned task processing method or automatic question-answering method.The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the task processing method and the automatic question-answering method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the task processing method or the automatic question-answering method described above. One embodiment of the present disclosure also provides a computer program. When executed on a computer, the computer program causes the computer to perform the steps of the task processing method or the automatic question-answering method described above. The above is a schematic diagram of a computer program according to this embodiment. It should be noted that the technical solution of this computer program is based on the same concept as the technical solutions of the task processing method and the automatic question-answering method described above. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the task processing method or the automatic question-answering method described above. The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. In addition, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous. The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals. It should be noted that for ease of description, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present disclosure are not limited by the order of the actions described, as certain steps may be performed in other orders or simultaneously according to the embodiments of the present disclosure. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the embodiments of the present disclosure.In the above embodiments, the description of each embodiment has its own emphasis. For portions not described in detail in a particular embodiment, reference should be made to the relevant descriptions of other embodiments. The preferred embodiments disclosed above are merely intended to illustrate the present disclosure. The alternative embodiments do not describe all details in detail, nor do they limit the invention to the specific implementations described. Obviously, many modifications and variations are possible based on the content of the embodiments disclosed. These embodiments are selected and described in detail in this disclosure to better explain the principles and practical applications of the embodiments of the present disclosure, thereby enabling those skilled in the art to better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

Claims 1. A task processing method, comprising: Obtain a sequence of feature blocks of target task data, where the sequence of feature blocks includes a plurality of feature blocks; in the case that the memory of the current service unit is insufficient, query a target service unit from service units other than the current service unit, where the target service unit includes available memory for processing the feature blocks; call the current service unit to process the first feature block among the plurality of feature blocks to obtain a first processing result, and call the target service unit to process the second feature block among the plurality of feature blocks to obtain a second processing result; generate a task processing result of the target task data according to the first processing result and the second processing result.

2. The method according to claim 1, wherein querying for a target service unit from service units other than the current service unit comprises: Determine memory requirement information according to the difference between the current available memory of the current service unit and the processing memory corresponding to the plurality of feature blocks. Filter out a plurality of candidate service units from a plurality of service units according to the unit memory usage table and the memory requirement information, where the candidate service units include available memory that meets the memory requirement information. Filter out a target service unit from the plurality of candidate service units.

3. The method according to claim 2, wherein screening out the target service unit from the multiple candidate service units comprises: Send a memory borrowing request to the candidate service units according to the priorities of the plurality of candidate service units. In the case of receiving borrowing confirmation information returned by the first candidate service unit, determine the first candidate service unit as the target service unit, where the first candidate service unit is any one of the plurality of candidate service units.

4. The method according to claim 2 or 3, further comprising, before screening out a plurality of candidate service units from a plurality of service units according to the unit memory usage table and the memory requirement information: Obtain heartbeat information of a plurality of service units, where the heartbeat information includes available memory information and memory interaction information of the service units; construct a unit memory usage table according to the available memory information and the memory interaction information of the plurality of service units.

5. The method according to claim 4, after obtaining the heartbeat information of multiple service units, further comprising: Parse the memory interaction information to determine the memory interaction relationship between the plurality of service units. Construct a memory interaction graph with the service units as nodes and the memory interaction relationship as edges. Adjust the feature blocks corresponding to the service units according to the memory interaction graph.

6. The method according to claim 5, wherein adjusting the feature block corresponding to the service unit according to the memory interaction diagram comprises: Filter out service units to be adjusted from the memory interaction graph, where the service units to be adjusted form a circular relationship structure in the memory interaction graph; adjust the corresponding feature blocks of the service units to be adjusted according to the memory interaction information of the service units to be adjusted.

7. The method according to any one of claims 1 to 6, wherein the calling the target service unit to process the second feature block among the plurality of feature blocks to obtain a second processing result includes: Look up a target memory address in the address mapping table of the target service unit according to the attribute information of the second feature block. Store the second feature block to the target service unit according to the target memory address, and call the target service unit to process the second feature block.

8. The method according to any one of claims 1 to 7, wherein generating a task processing result of the target task data according to the first processing result and the second processing result comprises: Aggregate the first processing result and the second processing result to obtain an aggregated processing result. Generate a task processing result of the target task data according to the aggregated processing result.

9. The method according to any one of claims 1 to 8, wherein the obtaining of the feature block sequence of the target task data comprises: Obtain target task data. Extract features from the target task data to obtain a task feature sequence, where the task feature sequence includes multiple task data features; Construct a feature block sequence according to the feature block division condition and the multiple task data features.

10. Before the method according to any one of claims 1 to 9, where the current service unit is called to process the first feature block among the multiple feature blocks to obtain a first processing result, and the target service unit is called to process the second feature block among the multiple feature blocks to obtain a second processing result, it further includes: Obtain the current available memory of the current service unit; Divide the multiple feature blocks according to the current available memory to determine a first feature block and a second feature block, where the current available memory is used to process the first feature block.

11. An automatic question answering method, comprising: Obtain a feature block sequence of the question to be answered, where the feature block sequence includes multiple feature blocks; when the memory of the current service unit is insufficient, query a target service unit from service units other than the current service unit, where the target service unit includes available memory for processing the feature blocks; call the current service unit to process the first feature block among the multiple feature blocks to obtain a first processing result, and call the target service unit to process the second feature block among the multiple feature blocks to obtain a second processing result; generate a reply result to the question to be answered according to the first processing result and the second processing result.

12. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10 or claim 11 are implemented.

13. A computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 or claim 11 are implemented.

14. A computer program, when the computer program is executed on a computer, causes the computer to execute the steps of the method according to any one of claims 1 to 10 or claim 11.

Citation Information

Patent Citations

  • Data processing method, device and distributed service system

    CN109343962A

  • Data distribution method and device, model training method and device of data distribution method and device, and computing cluster

    CN109710406A

  • FPGA-based neural network operation method, device and equipment

    CN111860810A

  • Method, device and system for deep learning model training

    CN114139723A

  • Method, device and equipment for realizing open domain questions and answers of large-scale language model

    CN116662509A