Task processing method and device, computer equipment, storage medium and program product
By using a shared encoder to encode and cache features during content recommendation and annotation, the problem of resource waste is solved, the calculation time and resources are significantly reduced, and the consistency and accuracy of the task model are improved.
Patent Information
- Application Number
- CN202510118259.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has wasted resources in the process of content recommendation and labeling. The same text content is repeatedly called and calculated, which increases the computing load and time overhead. The independent operation of multiple models occupies more memory and processor resources, reducing the overall performance and efficiency of the system.
The content is encoded through the shared encoder and cached the encoded content features. The features are read directly from the cache when responding to the task processing request, avoiding repeated encoding, so that different tasks can multiplex the results of the shared encoder.
Reduce repeated calculations, greatly reduce computing time and resource consumption, improve the consistency and accuracy of different task models, save training overhead, and improve the overall performance and efficiency of the system.
Smart Images

Figure CN120066713A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence, and particularly to a task processing method, apparatus, computer device, storage medium, and program product. Background Art
[0002] Artificial intelligence technology is applied to more and more production and life scenarios. For example, in the field of content recommendation, the tags to which the content belongs can be marked by a large language model, and then the content of interest to each user can be pushed to different users based on the tags.
[0003] In related technologies, the tags to which the content belongs can include tags in multiple vertical domains. For example, when the vertical domain is the content type, the tags can be "finance", "technology", "entertainment", and "culture"; when the vertical domain is advertisement recognition, the tags can be "contains advertisement" and "does not contain advertisement". For different tag annotation tasks in different vertical domains, different models are usually used in related technologies to extract features from the text content and perform tag annotation respectively.
[0004] However, there is a problem of resource waste in the above manner. On the one hand, the same text content is repeatedly called and calculated, increasing the computational load and time overhead; on the other hand, the independent operation of multiple models occupies more memory and processor resources, reducing the overall performance and efficiency of the computer device. Summary of the Invention
[0005] Embodiments of the present application provide a task processing method, apparatus, computer device, storage medium, and program product. The technical solutions are as follows:
[0006] On the one hand, embodiments of the present application provide a task processing method, the method including:
[0007] Encoding the content through a shared encoder and caching the content features obtained by encoding, where the shared encoder is an encoder shared by different task models, the different task models are used to process different tasks, and the shared encoder is obtained by joint training based on different tasks;
[0008] In response to a first task processing request for the first content, reading the first content feature corresponding to the first content from the cache;
[0009] Inputting the first content feature and a first task prompt into a first task model to obtain a first task processing result output by the first task model, where the first task model is used to process the first task, and the first task prompt is used to describe the task requirements of the first task through text, and different tasks correspond to different task prompts and different task models.
[0010] On the other hand, an embodiment of the present application provides a task processing device, which includes:
[0011] A cache module, configured to encode content through a shared encoder and cache the content features obtained by encoding. The shared encoder is an encoder shared by different task models, and different task models are used to process different tasks. The shared encoder is obtained by joint training based on different tasks;
[0012] A reading module, configured to read, in response to a first task processing request for a first piece of content, the first content feature corresponding to the first piece of content from the cache;
[0013] A processing module, configured to input the first content feature and a first task prompt into a first task model to obtain a first task processing result output by the first task model. Among them, the first task model is used to process the first task, and the first task prompt is used to describe the task requirements of the first task through text. Different tasks correspond to different task prompts and different task models.
[0014] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory. At least one computer instruction is stored in the memory, and the at least one computer instruction is loaded and executed by the processor to implement the method described in the above aspect.
[0015] On the other hand, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer instruction is stored, and the computer instruction is loaded and executed by a processor to implement the method described in the above aspect.
[0016] On the other hand, an embodiment of the present application provides a computer program product, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners of the above aspect.
[0017] In the model inference stage of the embodiment of the present application, since the computer device caches the content features obtained by encoding the content by the shared encoder, and the shared encoder is an encoder shared by different task models, when it is necessary to perform different task processing on the first piece of content, the first content feature corresponding to the first piece of content can be directly read from the cache without encoding the first piece of content again. That is, different tasks can reuse the first content feature obtained by encoding the first piece of content by the shared encoder. On the one hand, repeated calculations are avoided, and the calculation time and resource consumption are greatly reduced. On the other hand, the first content features shared by different tasks improve the consistency and accuracy of different task models. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 is a schematic diagram of implementing the label annotation task of multiple vertical domains in the related art;
[0020] Figure 2 is a flowchart of the task processing method provided by an exemplary embodiment of the present application;
[0021] Figure 3 is a schematic diagram of the jointly trained shared encoder based on different tasks provided by an exemplary embodiment of the present application;
[0022] Figure 4 is a schematic diagram of caching content features in the form of cache pairs provided by an exemplary embodiment of the present application;
[0023] Figure 5 is a schematic diagram of the training process of the dimensionality reduction matrix and the dimensionality increase matrix provided by an exemplary embodiment of the present application;
[0024] Figure 6 is a flowchart of the task processing method provided by an exemplary embodiment of the present application in the case of a new task in addition to each i-th task;
[0025] Figure 7 is a schematic diagram of implementing the task processing method through the LLaMa model provided by an exemplary embodiment of the present application;
[0026] Figure 8 is a comparison diagram of the effects of the related technology and the task processing method provided by the present application;
[0027] Figure 9 is a structural block diagram of the task processing device provided by an exemplary embodiment of the present application;
[0028] Figure 10 is a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.
[0030] In traditional technical solutions, to implement tagging of multi-vertical domain labels, multiple independent models are usually used to extract features and label the content respectively. When multi-vertical domain label annotation is required for the content, the computer device will sequentially call the models corresponding to these vertical domains to perform multiple calculations and analyses on the same content.
[0031] See Figure 1 , Figure 1 which is a schematic diagram of implementing the multi-vertical domain label annotation task in the related art.
[0032] Taking text content as an example, for vertical domain 1 (such as content type), in the related art, the text content, the rules and questions of task 1 (content type classification) corresponding to vertical domain 1 are input into model 1 corresponding to task 1 for feature extraction and content type classification, and model 1 outputs label 1 corresponding to task 1 (such as the content type is finance); for vertical domain 2 (such as whether it contains advertisements), in the related art, the text content, the rules and questions of task 2 (advertisement inclusion situation classification) corresponding to vertical domain 2 are input into model 2 corresponding to task 2 for feature extraction and advertisement inclusion situation classification, and model 2 outputs label 2 corresponding to task 2 (such as contains advertisements)... And so on, for different tasks, it is necessary to input the text content, the rules and prompts corresponding to the tasks into different task models for feature extraction and task processing to obtain the task processing results.
[0033] The above processing method may not have an obvious impact in small-scale applications, but in large-scale data processing and real-time application scenarios, the waste of resources will increase significantly, affecting the throughput and user experience of the system. For example, in a real-time recommendation system, it is necessary to quickly perform multi-label annotation on a large amount of content, and the traditional multi-model independent calculation method is difficult to meet the requirements of high efficiency. First, the same content is repeatedly called and calculated, increasing the computational load and time overhead of the system. Second, the independent operation of multiple models occupies more memory and processor resources, reducing the overall performance and efficiency of the system. In addition, the lack of cooperation between models may lead to repeated calculation of the same or similar feature extraction and analysis processes, and the existing calculation results cannot be fully utilized.
[0034] To seek a technical solution that can efficiently and accurately perform multi-task processing on content, avoid repeated calculation, and improve system performance, this application provides a task processing method.
[0035] See Figure 2 , Figure 2 which is a flowchart of the task processing method provided by an exemplary embodiment of this application. In some embodiments, this method is executed by a computer device (such as a server). This method includes the following steps.
[0036] Step 201: Encode the content through a shared encoder and cache the content features obtained by the encoding. The shared encoder is an encoder shared by different task models. The different task models are used to process different tasks, and the shared encoder is obtained through joint training based on different tasks.
[0037] Optionally, the content includes one or more of text content, picture content, video content, audio content, live broadcast content, or any other possible multimodal content, and there is no limitation on this.
[0038] Regarding the acquisition method of the content, only by way of example, the computer device can use the article content of each article in the news application as the text content, or use the photos in the social application as the picture content, or other any possible methods can also be used to acquire the content, and there is no limitation on this.
[0039] It should be noted that in the process of collecting relevant data of the user (such as the text content, picture content, etc. published by the user), this application can display a prompt interface, a pop-up window, or output a voice prompt message. The prompt interface, pop-up window, or voice prompt message is used to prompt the user that their relevant data is being collected currently, so that this application only starts to execute the relevant steps of acquiring the user's relevant data after obtaining the confirmation operation of the user on the prompt interface or the pop-up window. Otherwise (that is, when the confirmation operation of the user on the prompt interface or the pop-up window is not obtained), the relevant steps of acquiring the user's relevant data are ended, that is, the relevant data of the user is not acquired. In other words, the information (including but not limited to the user device information, user personal information, real-time location of the user), data (including but not limited to the data for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the text content, picture content, etc. published by the user involved in this application are all obtained under sufficient authorization.
[0040] The different task models are used to process different tasks. Only by way of example, the different tasks are content classification tasks for different vertical domains. Task 1 is a content classification task for vertical domain 1 (such as content type), and task model 1 is used to label the content type to which the content belongs; Task 2 is a content classification task for vertical domain 2 (such as whether it contains advertisements), and task model 2 is used to label the advertisement inclusion situation of the content...
[0041] Only by way of example, the different tasks can also be tasks such as negative energy recognition task for text content, news clickbait recognition task, induced like-cheating recognition task, timeliness analysis task, etc. Through different tasks, multi-angle analysis of the content can be realized.
[0042] The shared encoder is an encoder shared by different task models, that is, different task models all perform task processing on the content features obtained by encoding the content using the shared encoder.
[0043] Regarding the specific model structure of the shared encoder, in some embodiments, the shared encoder can be a large language model (LLM), or the shared encoder can include the model structure in the large language model.
[0044] By way of example only, the shared encoder can be the LLaMa model (Large Language Model Meta AI), or the shared encoder can be a partial model structure in the LLaMa model (such as the first 31 transformer blocks therein).
[0045] In some embodiments, the shared encoder is obtained by joint training based on different tasks, that is, the shared encoder can be obtained by joint training with different task models corresponding to different tasks.
[0046] See Figure 3 , Figure 3 is a schematic diagram showing that the shared encoder provided by an exemplary embodiment of this application is obtained by joint training based on different tasks.
[0047] In some embodiments, the computer device encodes the sample content through the shared encoder 301 to obtain sample content features.
[0048] In some embodiments, the computer device inputs the sample content features into different task models for processing different tasks to obtain sample task processing results.
[0049] For example, input the sample content features into the task model 310-1 for processing task 1 to obtain the sample first task processing result; input the sample content features into the task model 310-2 for processing task 2 to obtain the sample second task processing result;... input the sample content features into the task model 310-n for processing task n to obtain the sample nth task processing result.
[0050] In some embodiments, the computer device updates the model parameters of the shared encoder 301 and each task model according to the difference between each sample processing result and the corresponding task result truth value.
[0051] For example, the first loss is determined based on the difference between the processing result of the first task for the sample and the true value of the first task result, the second loss is determined based on the difference between the processing result of the second task for the sample and the true value of the second task result, ……, the nth loss is determined based on the difference between the processing result of the nth task for the sample and the true value of the nth task result, and the sum of the first loss to the nth loss is used as the total loss to jointly update the model parameters of the shared encoder 301 and the task models 310-1, 310-2, ……, 310-n.
[0052] In some embodiments, the computer device can pre-encode each different piece of content through the shared encoder and cache the respective content features corresponding to each different piece of content. Among them, the cache (Cache) for storing each content feature is a memory that can perform high-speed data exchange.
[0053] Optionally, the content feature obtained by encoding the content by the shared encoder can be in the form of a feature vector or a feature matrix. By way of example only, the text content feature obtained by encoding the text content can be a feature vector. By way of example only, the picture content feature obtained by encoding the picture content can be in the form of a feature map, and the feature value corresponding to each pixel point on the feature map is the value obtained by encoding the pixel point by the shared encoder.
[0054] In a possible implementation manner, the computer device can store the content features in the form of cache pairs. Among them, the cache index in the cache pair is the content identifier corresponding to the content, and the cache value in the cache pair is the content feature obtained by encoding.
[0055] Step 202, in response to a first task processing request for the first content, read the first content feature corresponding to the first content from the cache.
[0056] The first task processing request is used to request the processing of the first task. For example, when the first content is article A and the first task is a content type classification task, the first task processing request is used to request to determine the content type classification to which article A belongs.
[0057] In a possible scenario, when the computer device has pre-cached the first content feature corresponding to the first content, the computer device searches for the first content feature corresponding to the first content from the respective content features in the cache and reads the first content feature.
[0058] In another possible scenario, when the computer device fails to find the first content feature corresponding to the first content, the computer device can extract the first content feature of the first content through the shared encoder and write the first content feature into the cache for subsequent use in other tasks for the first content.
[0059] Step 203: Input the first content feature and the first task prompt into the first task model to obtain the first task processing result output by the first task model. Here, the first task model is used to process the first task, and the first task prompt is used to describe the task requirements of the first task in text. Different tasks correspond to different task prompts and different task models.
[0060] In some embodiments, the task model can be a large language model (LLM), or the task model can include the model structure in the large language model.
[0061] By way of example only, the task model can be the LLaMa model (Large Language Model Meta AI), or the task model can be a partial model structure in the LLaMa model (such as the 32nd transformer block therein).
[0062] The task prompt is used to describe the task requirements of the task in text. Optionally, the task prompt is used to describe the rules and questions of the task.
[0063] By way of example only, if the first task is a content type classification task, the first task prompt can be "Please determine which content type the given content belongs to according to the provided content? For example, finance, news, culture, or entertainment, etc."
[0064] Optionally, the different task prompts corresponding to different tasks can be determined by being pre-set in the computer device.
[0065] In some embodiments, the different task models corresponding to different tasks can be pre-trained. For example, the Figure 3 method can be used to jointly train each task model with a shared encoder, or alternatively, each task model can be trained separately, and there is no limitation on this.
[0066] In some embodiments, the computer device inputs the first content feature and the first task prompt into the first task model to obtain the first task processing result output by the first task model.
[0067] By way of example only, the computer device inputs the feature vector (the first content feature) of article A (the first content) read from the cache into the content type classification model (the first task model) to obtain that the content type of article A is the finance type (the first task processing result).
[0068] In summary, in the model inference stage, since the computer device caches the content features encoded by the shared encoder for the content, and the shared encoder is shared by different task models, when different tasks need to be performed on the first content, the first content features corresponding to the first content can be directly read from the cache without encoding the first content again. That is, different tasks can reuse the first content features encoded by the shared encoder for the first content. On the one hand, this avoids repeated calculations and significantly reduces the computing time and resource consumption. On the other hand, the first content features shared by different tasks improve the consistency and accuracy of different task models.
[0069] In the model training stage, the shared encoder is jointly trained based on different tasks. Compared with the related art where separate encoders need to be trained for different tasks, only one encoder shared by different task models needs to be trained. Therefore, the training overhead can also be saved, and the overall performance and efficiency of the computer device can be improved.
[0070] When caching the content features corresponding to the content, the computer device can use the form of cache pairs for caching.
[0071] See Figure 4 , Figure 4 which is a schematic diagram of caching content features in the form of cache pairs provided by an exemplary embodiment of the present application.
[0072] In some embodiments, the computer device encodes the content through the shared encoder and constructs cache pairs.
[0073] Among them, the cache index in the cache pair is the content identifier corresponding to the content, and the cache value in the cache pair is the content feature obtained by encoding.
[0074] In some embodiments, the shared encoder 410 encodes the content 401 to obtain the content feature 402.
[0075] When the shared encoder is Encoder and the content is x, the content feature Z can be expressed by the following formula.
[0076] Z = Encoder(x; θ E );
[0077] where Encoder is a pre-trained shared encoder for extracting the content features of the content. θ E is the model parameter of the shared encoder, which has been learned through pre-training. The content feature Z is usually a high-dimensional representation containing global semantic information, such as a vector or tensor of a fixed size.
[0078] In some embodiments, the computer device determines the content identifier 402 corresponding to the content 401.
[0079] Among them, the content identifier is used to uniquely identify the corresponding content, and different contents correspond to different content identifiers.
[0080] Optionally, the computer device uses the hash operation result of the content as the content identifier corresponding to the content.
[0081] When the content is x and the hash function is Hash, the content identifier k can be expressed by the following formula.
[0082] k = Hash(x);
[0083] Optionally, the unique marking symbol corresponding to the content is determined as the content identifier.
[0084] Merely by way of example, the computer device can sort each content and determine a sorting number (ID), and determine the sorting number as the content identifier.
[0085] In some embodiments, the computer device constructs a cache pair (k, Z) with the content identifier 403 as the cache index and the content feature 402 as the cache value. Wherein, k is the content identifier and Z is the content feature.
[0086] In some embodiments, the computer device writes the cache pair into the cache.
[0087] In some embodiments, in response to a first task processing request for the first content, the computer device searches for a first cache pair with the cache index being the first content identifier from the cache pairs in the cache, and determines the first content feature based on the cache value in the first cache pair.
[0088] Among them, the first content identifier is the content identifier corresponding to the first content.
[0089] Merely by way of example, when the first content identifier of the first content is "00001", the computer device searches for a first cache pair with the cache index being "00001" from the cache pairs in the cache, and uses the cache value of the first cache pair as the first content feature.
[0090] The data volume of the content feature is usually large. For example, when the content feature is a long sequence, the cache pair may become unable to be efficiently stored and transmitted. For this reason, in some embodiments, the content feature can also be compressed and stored, and restored to the original data volume when read, so as to achieve the effect of saving memory and improving the writing and reading efficiency.
[0091] In some embodiments, the computer device compresses the content feature encoded by the shared encoder for the content to obtain a compressed content feature.
[0092] Among them, the feature dimension of the compressed content feature is smaller than the feature dimension of the content feature.
[0093] In a possible scenario regarding the specific manner of compressing content features, a computer device determines the compressed content features as the product of the content features encoded by a shared encoder and a dimensionality reduction matrix.
[0094] Using hidden states to represent the content features and W kv to represent the dimensionality reduction matrix, the compressed content features compressed kv can be expressed by the following formula:
[0095] compressed kv = W kv · hidden states ;
[0096] Optionally, the dimensionality reduction matrix can be the model parameters of a linear layer with low-rank factorization.
[0097] In some embodiments, a computer device constructs cache pairs and writes the cache pairs into the cache. Among them, the cache index in the cache pair is the content identifier, and the cache value in the cache pair is the compressed content features.
[0098] In some embodiments, a computer device reads the first compressed content features in the first cache pair and restores the feature dimension of the first compressed content features to obtain the first content features.
[0099] Regarding the specific manner of restoring the feature dimension, in a possible scenario, a computer device uses the product of the first compressed content features and a dimensionality increase matrix as the first content features.
[0100] Optionally, a computer device can first perform normalization processing on the first compressed content features, and then use the product of the result after normalization processing and the dimensionality increase matrix as the first content features.
[0101] Using compressed kv to represent the compressed content features, W kvb to represent the dimensionality increase matrix, and LayerNorm to represent the normalization layer, the first content features kv after dimension restoration can be expressed by the following formula:
[0102] kv = W kvb · LayerNorm(compressed kv );
[0103] Among them, the dimensionality reduction matrix and the dimensionality increase matrix can be matrices obtained through pre-training.
[0104] See Figure 5 , Figure 5It is a schematic diagram of the training process of the dimensionality reduction matrix and the dimensionality increase matrix provided by an exemplary embodiment of the present application.
[0105] In some embodiments, the computer device encodes the sample content through the shared encoder 510 to obtain the sample content features.
[0106] In some embodiments, the computer device uses the product of the dimensionality reduction matrix 521 and the sample content features as the sample compressed content features.
[0107] In some embodiments, the computer device determines the product of the dimensionality increase matrix 522 and the sample compressed content features as the sample restored content features.
[0108] In some embodiments, the computer device inputs the sample restored content features and the i-th task prompt into the i-th task model to obtain the sample i-th task processing result output by the i-th task model.
[0109] Wherein, the i-th task model is used to process the i-th task, the i-th task prompt is used to describe the task requirements of the i-th task through text, and i is a positive integer. Optionally, i is any positive integer between 1 and n. N is the number of task models.
[0110] For example only, the computer device inputs the sample restored content features and the first task prompt into the first task model 530-1 to obtain the sample first task processing result; inputs the sample restored content features and the second task prompt into the second task model 530-2 to obtain the sample second task processing result;... inputs the sample restored content features and the n-th task prompt into the n-th task model 530-n to obtain the sample n-th task processing result.
[0111] In some embodiments, the computer device updates the matrix parameters of the dimensionality reduction matrix and the dimensionality increase matrix according to the difference between the sample i-th task processing result and the i-th task result true value.
[0112] For example, determine the first loss according to the difference between the sample first task processing result and the first task result true value, determine the second loss according to the difference between the sample second task processing result and the second task result true value,..., determine the n-th loss according to the difference between the sample n-th task processing result and the n-th task result true value, and use the sum of the first loss to the n-th loss as the total loss to update the matrix parameters of the dimensionality reduction matrix 521 and the dimensionality increase matrix 522.
[0113] In one possible implementation, the computer device can separately train the dimensionality reduction matrix and the dimensionality increase matrix when the shared encoder and each task model are trained; in another possible implementation, the computer device can also train the dimensionality reduction matrix and the dimensionality increase matrix simultaneously during the joint training process of the shared encoder and each task model, and there is no limitation on this.
[0114] In this embodiment, the dimensionality reduction matrix obtained through training can write the compressed content features into the cache, thereby reducing memory usage and improving the writing efficiency; when reading the compressed content features from the cache, the dimensionality increase matrix obtained through training can restore the feature dimension of the compressed content features. When the dimensionality increase matrix is well-trained, the performance of the model will not be significantly affected.
[0115] Regarding the training method of the shared encoder, in a possible scenario, the shared encoder and each task model are jointly trained.
[0116] In some embodiments, the computer device encodes the sample content through the shared encoder to obtain the sample content features.
[0117] In some embodiments, the computer device inputs the sample content features and the i-th task prompt into the i-th task model to obtain the sample i-th task processing result output by the i-th task model.
[0118] Among them, the i-th task model is used to process the i-th task, and the i-th task prompt is used to describe the task requirements of the i-th task through text, where i is a positive integer.
[0119] In some embodiments, the computer device updates the model parameters of the shared encoder and each i-th task model according to the difference between the sample i-th processing result and the true value of the i-th task result.
[0120] In the actual application process of task processing, new tasks may be generated.
[0121] For example only, Task 1 is a content classification task in vertical domain 1 (such as content type), and Task 2 is a content classification task in vertical domain 2 (such as whether it contains advertisements). In the case of a new content classification requirement, the new Task 3 can be a content classification task in vertical domain 3 (such as whether there is an induced like-cheating behavior).
[0122] See Figure 6 , Figure 6 is a flowchart of a task processing method provided by an exemplary embodiment of the present application in the case of new tasks in addition to each i-th task. The process includes the following steps.
[0123] Step 610, in the case of new tasks in addition to each i-th task, the computer device trains a new task model based on the sample content.
[0124] Among them, the parameters of the shared encoder are fixed during the training process of the new task model.
[0125] For example, a computer device can collect sample content with the true value of the new task result, input the sample content into the shared encoder to obtain the sample content features, then input the sample content features into the new task model to obtain the new task processing result, construct a new loss based on the difference between the new task processing result and the true value of the new task result, and update the model parameters of the new task model based on the new loss.
[0126] Step 620, determine the model performance of the new task model through the test sample set.
[0127] Optionally, the sample content in the test sample set is different from the sample content in the training sample set.
[0128] Optionally, the metrics for evaluating the model performance include accuracy, which is used to describe the proportion of the number of samples correctly predicted by the model to the total number of samples.
[0129] Optionally, the metrics for evaluating the model performance include recall, which is used to measure the proportion of positive examples correctly identified by the model among all samples that are actually positive examples.
[0130] In addition, those skilled in the art can also determine the metrics for evaluating the model performance according to actual needs, and there is no limitation on this.
[0131] Step 630, determine whether the model performance meets the performance requirements.
[0132] Optionally, the performance requirements are the requirements preset by the computer device. For example, the accuracy reaches 95% or the recall reaches 80%, and there is no limitation on this.
[0133] In the case where the model performance does not meet the performance requirements, execute step 641; in the case where the model performance meets the performance requirements, execute step 642.
[0134] Step 641, update the model parameters of the shared encoder and the new task model based on the training sample set so that the model performance meets the performance requirements.
[0135] In the case where the model performance of the new task model does not meet the performance requirements, it indicates that the shared encoder fails to accurately extract the content features required by the new task model. Therefore, at this time, the shared encoder needs to be retrained, rather than simply using the shared encoder jointly trained with the previous task models.
[0136] Among them, the training sample set contains sample content with the true value of the new task result. The computer device inputs the sample content in the training sample set into the shared encoder to obtain sample content features; then inputs the sample content features into the new task model to obtain the new task processing result; constructs a new loss based on the difference between the new task processing result and the true value of the new task result, and updates the model parameters of the shared encoder based on the new loss, or updates the model parameters of the shared encoder and the new task model, so as to train a shared encoder more suitable for the new task model and improve the model performance.
[0137] Step 642, in response to a new task processing request for the first content, input the first content feature read from the cache and the new task prompt into the new task model to obtain the new task processing result output by the new task model.
[0138] By way of example only, the computer device inputs the feature vector (the first content feature) of article A (the first content) read from the cache and the new task prompt into the new task model to obtain the classification result (the new task processing result) that article A has the risk of inducing false likes.
[0139] In this embodiment, in the case of a new task, if the new task is relatively close to each previous task, the shared encoder can obtain better new task processing results with the new task model, and the previously trained shared encoder can be directly reused, reducing the model training cost; if the new task is quite different from each previous task, resulting in poor model performance of the shared encoder and the new task model, it is necessary to retrain the shared encoder based on the sample content with the true value of the new task result to improve the feature encoding ability of the shared encoder for the new task.
[0140] Regarding the specific structures of the shared encoder and the task model, in some embodiments, the shared encoder is the encoding network in a large language model, and the task model is the task network in a large language model.
[0141] In some embodiments, the shared encoder includes a first attention network.
[0142] By way of example only, the shared encoder is the first 31 transformer blocks in the LLaMa model, where the first 31 transformer blocks all belong to the first attention network.
[0143] In some embodiments, the task model includes a second attention network and a task processing network.
[0144] For example, the task model is the 32nd transformer block in the LLaMa model. Among them, the 32nd transformer block belongs to the second attention network.
[0145] For example, the task processing network is the remaining model structure after the 32nd transformer block in the LLaMa model (for example, including the RMSNorm layer, the Linear layer, and the Softmax layer).
[0146] In some embodiments, the computer device encodes the content through the encoding network in the large language model and caches the content features obtained by encoding.
[0147] In some embodiments, the computer device inputs the first content feature and the first task prompt into the task network in the large language model to obtain the first task processing result.
[0148] Optionally, the computer device encodes the content through the first attention network to obtain the first K matrix and the first V matrix, and caches the first K matrix and the first V matrix as content features.
[0149] See Figure 7 , Figure 7 is a schematic diagram of a task processing method implemented through the LLaMa model provided by an exemplary embodiment of the present application.
[0150] The shared encoder 710 includes the embedding layer (input embedding) in the LLaMa model and the first 31 transformer blocks (including Figure 7 Block1, Block2,... Block31 in
[0151] Among them, the first 31 transformer blocks all belong to the first attention network. The first task model 720 includes the second attention network and the task processing network. Among them, the 32nd transformer block in the LLaMa model belongs to the second attention network.
[0152] Regarding the specific calculation process of the second attention network (Block 32), as shown in Figure 7 the enlarged view of Block 32 in
[0153] By way of example only, the second K matrix includes Key token 5, Key token 7, and Key token 7.
[0154] By way of example only, the second V matrix includes Value token 5, Value token 7, and Value token 7.
[0155] In some embodiments, the computer device performs operations on the K matrix obtained by concatenating the first K matrix and the second K matrix, the V matrix obtained by concatenating the first V matrix and the second V matrix, and the Q matrix through the second attention network to obtain attention values.
[0156] By way of example only, if the first K matrix includes Key token 1 to 4 and the second K matrix includes Key token 5 to 7, then the concatenated K matrix includes Key token 1 to 7.
[0157] By way of example only, if the first V matrix includes Value token 1 to 4 and the second V matrix includes Value token 5 to 7, then the concatenated V matrix includes Value token 1 to 7.
[0158] In some embodiments, the computer device performs operations on the Q matrix, the K matrix, and the V matrix through the second attention network to obtain the attention value Attention(Q, K, V).
[0159] In some embodiments, the computer device inputs the attention value into the task processing network to obtain a first task processing result.
[0160] See Figure 8 , Figure 8 is a diagram comparing the effects of the related art and the task processing method provided by this application.
[0161] As Figure 8 shown, in the text content classification scenarios of different vertical domains, the effects of the related art and the task processing method provided by this application are respectively compared on tasks 1 to 4. Among them, the precision-recall rate of the task processing method provided by this application has been improved on each task.
[0162] In addition, since different amounts of computation will result in different computational costs, when completing the same task, the computational cost of using the related art is higher than that of using the task processing method provided by this application. Therefore, this application can greatly reduce the amount of computation and save computational costs.
[0163] In some embodiments, the content is text content, the task model is a text content classification model, and different task models are used to process text content classification tasks in different vertical domains.
[0164] In some embodiments, the computer device inputs the first text content feature and the text content classification hint of the first vertical domain into the first text content classification model to obtain the first text content classification result.
[0165] Optionally, the text content classification model may be a labeling model of classification labels.
[0166] Optionally, the text content classification model may be a binary classification model or a multi-classification model, and there is no limitation on this.
[0167] The task processing method proposed in this application can be applied to at least the following possible scenarios.
[0168] (1) Text content classification scenarios for different vertical domains.
[0169] In this scenario, different tasks are content classification tasks for different vertical domains.
[0170] For example, Task 1 is a content classification task for vertical domain 1 (such as content type), and Task Model 1 is used to label the content type to which the content belongs; Task 2 is a content classification task for vertical domain 2 (such as whether it contains advertisements), and Task Model 2 is used to label the advertisement inclusion situation of the content; Task 3 is a content classification task for vertical domain 3 (such as whether there is an induced like-cheating behavior), and Task Model 3 is used to label whether there is an induced like-cheating behavior in the content.
[0171] By using the task processing method provided in this application, read the text content feature obtained by encoding the text content by the shared encoder from the cache, and input the text content feature and the i-th task hint into the i-th task model, then the i-th task processing result can be obtained, where i is 1, 2, or 3.
[0172] (2) Image content segmentation scenarios for different segmentation objects.
[0173] In this scenario, different tasks are image content segmentation tasks for different segmentation objects.
[0174] For example, Task 1 is an image content segmentation task for segmenting people in the image, Task 2 is an image content segmentation task for segmenting animals in the image, and Task 3 is an image content segmentation task for segmenting plants in the image.
[0175] By using the task processing method provided in this application, read the image content feature obtained by encoding the image content by the shared encoder from the cache, and input the image content feature and the i-th task hint into the i-th task model, then the i-th task processing result can be obtained, where i is 1, 2, or 3.
[0176] Text content classification scenarios for different vertical domains
[0177] In an alternative example, the task processing method provided in the embodiments of the present application is introduced by taking the text content classification scenario in different vertical domains as an example.
[0178] In a possible scenario, the task processing method can be executed by a server.
[0179] In some embodiments, the text content is encoded by a shared encoder, and the text content features obtained by encoding are cached.
[0180] Among them, the shared encoder is an encoder shared by different text content classification models. The different text content classification models are used to process text content classification tasks in different vertical domains, and the shared encoder is jointly trained based on text content classification tasks in different vertical domains;
[0181] In some embodiments, in response to a first task processing request for the first text content, the first text content feature corresponding to the first text content is read from the cache.
[0182] In some embodiments, the first text content feature and the first task prompt are input into the first text content classification model to obtain a first task processing result output by the first text content classification model. Among them, the first text content classification model is used to process the first task, and the first task prompt is used to describe the task requirements of the first task through text. Different tasks correspond to different task prompts and different text content classification models.
[0183] Optionally, the text content is encoded by a shared encoder, and a cache pair is constructed. The cache index in the cache pair is the text content identifier corresponding to the text content, and the cache value in the cache pair is the text content feature obtained by encoding.
[0184] Optionally, the cache pair is written into the cache.
[0185] Optionally, in response to a first task processing request for the first text content, a first cache pair with a cache index being the first text content identifier is searched for in the cache pairs in the cache. The first text content identifier is the text content identifier corresponding to the first text content; based on the cache value in the first cache pair, the first text content feature is determined.
[0186] Optionally, the hash operation result of the text content is used as the text content identifier corresponding to the text content; or, the unique marking symbol corresponding to the text content is determined as the text content identifier.
[0187] Optionally, the text content features obtained by encoding the text content by the shared encoder are compressed to obtain compressed text content features, and the feature dimension of the compressed text content features is smaller than the feature dimension of the text content features.
[0188] Optionally, a cache pair is constructed, where the cache index in the cache pair is the text content identifier, and the cache value in the cache pair is the compressed text content feature.
[0189] Optionally, the product of the text content feature obtained by encoding the text content by the shared encoder and the dimensionality reduction matrix is determined as the compressed text content feature.
[0190] Optionally, the first compressed text content feature in the first cache pair is read.
[0191] Optionally, the feature dimension of the first compressed text content feature is restored to obtain the first text content feature.
[0192] Optionally, the product of the first compressed text content feature and the dimensionality increase matrix is used as the first text content feature.
[0193] See Figure 9 , Figure 9 which is a structural block diagram of a task processing device provided by an exemplary embodiment of the present application. The device includes:
[0194] A cache module 901, configured to encode content through a shared encoder and cache the encoded content feature. The shared encoder is an encoder shared by different task models. The different task models are used to process different tasks, and the shared encoder is jointly trained based on different tasks;
[0195] A reading module 902, configured to read, in response to a first task processing request for a first content, a first content feature corresponding to the first content from the cache;
[0196] A processing module 903, configured to input the first content feature and a first task prompt into a first task model to obtain a first task processing result output by the first task model. Wherein, the first task model is used to process a first task, and the first task prompt is used to describe the task requirements of the first task through text. Different tasks correspond to different task prompts and different task models.
[0197] Optionally, the cache module 901 is configured to:
[0198] Encode the content through the shared encoder and construct a cache pair. The cache index in the cache pair is the content identifier corresponding to the content, and the cache value in the cache pair is the encoded content feature;
[0199] Write the cache pair into the cache;
[0200] Optionally, the reading module 902 is configured to:
[0201] In response to the first task processing request for the first content, look up the first cache pair with a cache index of the first content identifier in the cache pairs in the cache, where the first content identifier is the content identifier corresponding to the first content;
[0202] Based on the cache value in the first cache pair, determine the first content feature.
[0203] Optionally, the cache module 901 is configured to:
[0204] Use the hash operation result of the content as the content identifier corresponding to the content; or,
[0205] Determine the unique marking symbol corresponding to the content as the content identifier.
[0206] Optionally, the cache module 901 is configured to:
[0207] Compress the content feature encoded by the shared encoder for the content to obtain a compressed content feature, where the feature dimension of the compressed content feature is smaller than the feature dimension of the content feature;
[0208] Construct the cache pair, where the cache index in the cache pair is the content identifier, and the cache value in the cache pair is the compressed content feature.
[0209] Optionally, the reading module 902 is configured to:
[0210] Read the first compressed content feature in the first cache pair;
[0211] Restore the feature dimension of the first compressed content feature to obtain the first content feature.
[0212] Optionally, the cache module 901 is configured to:
[0213] Determine the product of the content feature encoded by the shared encoder for the content and the dimensionality reduction matrix as the compressed content feature;
[0214] Optionally, the reading module 902 is configured to:
[0215] Use the product of the first compressed content feature and the dimensionality increase matrix as the first content feature.
[0216] Optionally, the device further includes a training module, configured to:
[0217] Encode the sample content through the shared encoder to obtain a sample content feature;
[0218] Use the product of the dimensionality reduction matrix and the sample content feature as the sample compressed content feature;
[0219] Determine the product of the dimensionality-increasing matrix and the content feature of the compressed sample as the content feature after sample restoration;
[0220] Input the content feature after sample restoration and the i-th task prompt into the i-th task model to obtain the processing result of the i-th task of the sample output by the i-th task model, where the i-th task model is used to process the i-th task, and the i-th task prompt is used to describe the task requirements of the i-th task through text, and i is a positive integer;
[0221] Update the matrix parameters of the dimensionality-reducing matrix and the dimensionality-increasing matrix according to the difference between the processing result of the i-th task of the sample and the true value of the i-th task result.
[0222] Optionally, the training module is further configured to:
[0223] Encode the sample content through the shared encoder to obtain the content feature of the sample;
[0224] Input the content feature of the sample and the i-th task prompt into the i-th task model to obtain the processing result of the i-th task of the sample output by the i-th task model, where the i-th task model is used to process the i-th task, and the i-th task prompt is used to describe the task requirements of the i-th task through text, and i is a positive integer;
[0225] Update the model parameters of the shared encoder and each of the i-th task models according to the difference between the i-th processing result of the sample and the true value of the i-th task result.
[0226] Optionally, the training module is further configured to:
[0227] In the case of new tasks other than each of the i-th tasks, train a new task model based on the sample content, and the parameters of the shared encoder are fixed during the training process of the new task model;
[0228] Determine the model performance of the new task model through a test sample set;
[0229] In the case that the model performance does not meet the performance requirements, update the model parameters of the shared encoder and the new task model based on the training sample set so that the model performance meets the performance requirements;
[0230] Optionally, the processing module 903 is further configured to:
[0231] When the performance of the model meets the performance requirements, in response to a new task processing request for the first content, the first content feature read from the cache and the new task prompt are input into the new task model, and the new task processing result output by the new task model is obtained.
[0232] Optionally, the shared encoder includes a first attention network, and the cache module 901 is used for:
[0233] Encoding the content through the first attention network to obtain a first K matrix and a first V matrix;
[0234] Caching the first K matrix and the first V matrix as the content features;
[0235] Optionally, the processing module 903 is used for:
[0236] Inputting the first K matrix, the first V matrix, and the first task prompt encoded by the first attention network for the first content into the first task model to obtain the first task processing result.
[0237] Optionally, the task model includes a second attention network and a task processing network, and the processing module 903 is used for:
[0238] Encoding the first task prompt through the second attention network to obtain a second K matrix, a second V matrix, and a Q matrix;
[0239] Performing an operation on the K matrix obtained by splicing the first K matrix and the second K matrix, the V matrix obtained by splicing the first V matrix and the second V matrix, and the Q matrix through the second attention network to obtain an attention value;
[0240] Inputting the attention value into the task processing network to obtain the first task processing result.
[0241] Optionally, the shared encoder is the encoding network in a large language model, and the task model is the task network in the large language model;
[0242] Optionally, the cache module 901 is used for:
[0243] Encoding the content through the encoding network in the large language model and caching the content features obtained by encoding;
[0244] Optionally, the processing module 903 is used for:
[0245] Inputting the first content feature and the first task prompt into the task network in the large language model to obtain the first task processing result.
[0246] Optionally, the content is text content, and the task model is a text content classification model. Different task models are used to process text content classification tasks in different vertical domains;
[0247] Optionally, the processing module 903 is configured to:
[0248] Input the first text content feature and the text content classification hint of the first vertical domain into the first text content classification model to obtain the first text content classification result.
[0249] See Figure 10 , Figure 10 is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Specifically: The computer device 1000 includes a central processing unit (CPU) 1001, a system memory 1004 including a random access memory 1002 and a read-only memory 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The computer device 1000 further includes a basic input / output system (Input / Output, I / O system) 1006 for facilitating the transmission of information between various components within the computer, and a mass storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1015.
[0250] The basic input / output system 1006 includes a display 1008 for displaying information and input devices 1009 such as a mouse and a keyboard for user input of information. Among them, both the display 1008 and the input devices 1009 are connected to the central processing unit 1001 through an input / output controller 1010 connected to the system bus 1005. The basic input / output system 1006 may further include an input / output controller 1010 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, a printer, or other types of output devices.
[0251] The mass storage device 1007 is connected to the central processing unit 1001 through a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable medium provide non-volatile storage for the computer device 1000. That is to say, the mass storage device 1007 may include a computer-readable medium (not shown) such as a hard disk or a drive.
[0252] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several types. The above system memory 1004 and mass storage device 1007 can be collectively referred to as memory.
[0253] The memory stores one or more programs, the one or more programs are configured to be executed by one or more central processing units 1001, the one or more programs include computer instructions for implementing the above method, and the central processing unit 1001 executes the one or more programs to implement the methods provided by the above various method embodiments.
[0254] According to various embodiments of the present application, the computer device 1000 may also run by connecting to a remote computer on the network through a network such as the Internet. That is, the computer device 1000 may be connected to the network 1012 through the network interface unit 1011 connected to the system bus 1005, or in other words, the network interface unit 1011 may also be used to connect to other types of networks or remote computer systems (not shown).
[0255] The memory further includes one or more programs, the one or more programs are stored in the memory, and the one or more programs include steps for performing the methods executed by the computer device in the embodiments provided by the present application.
[0256] An embodiment of the present application further provides a computer-readable storage medium, in which at least one computer instruction is stored, and the at least one computer instruction is loaded and executed by a processor to implement the method described in the above embodiment. Optionally, the computer-readable storage medium may include: ROM, RAM, solid state drives (SSDs), optical discs, etc. Among them, RAM may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).
[0257] An embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners of the above aspects.
[0258] The foregoing are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A task processing method, characterized in that: The method comprises: Encode the content through a shared encoder and cache the content features obtained by encoding. The shared encoder is an encoder shared by different task models. Different task models are used to process different tasks. The shared encoder is obtained based on joint training of different tasks. In response to a first task processing request for first content, reading a first content feature corresponding to the first content from a cache; The first content feature and the first task prompt are input into a first task model to obtain a first task processing result output by the first task model, wherein the first task model is used to process a first task, the first task prompt is used to describe the task requirements of the first task through text, and different tasks correspond to different task prompts and different task models.
2. The method according to claim 1, characterized in that The encoding of the content by the shared encoder and caching of the encoded content features include: Encode the content by using the shared encoder, and construct a cache pair, wherein the cache index in the cache pair is a content identifier corresponding to the content, and the cache value in the cache pair is the content feature obtained by encoding; writing the cache pair into the cache; The step of reading, in response to the first task processing request for the first content, a first content feature corresponding to the first content from the cache includes: In response to the first task processing request for the first content, searching the cache pairs in the cache for a first cache pair whose cache index is a first content identifier, where the first content identifier is the content identifier corresponding to the first content; The first content characteristic is determined based on a cache value in the first cache pair.
3. The method according to claim 2, characterized in that The method further comprises: Using the hash operation result of the content as the content identifier corresponding to the content; or The unique marking symbol corresponding to the content is determined as the content identifier.
4. The method according to claim 2, characterized in that: The step of encoding the content by the shared encoder and constructing a cache pair includes: Compressing the content feature obtained by encoding the content by the shared encoder to obtain a compressed content feature, wherein the feature dimension of the compressed content feature is smaller than the feature dimension of the content feature; Constructing the cache pair, wherein the cache index in the cache pair is the content identifier, and the cache value in the cache pair is the compressed content feature; The determining the first content feature based on the cache value in the first cache pair includes: reading a first compressed content feature in the first cache pair; The feature dimension of the first compressed content feature is restored to obtain the first content feature.
5. The method according to claim 4, characterized in that The compressing the content feature obtained by encoding the content by the shared encoder to obtain the compressed content feature includes: Determine the compressed content feature by multiplying the content feature obtained by encoding the content by the shared encoder and the dimension reduction matrix; The restoring the feature dimension of the first compressed content feature to obtain the first content feature includes: The product of the first compressed content feature and the dimension-increased matrix is used as the first content feature.
6. The method according to claim 5, characterized in that The method further comprises: Encoding the sample content by the shared encoder to obtain sample content features; The product of the dimension reduction matrix and the sample content feature is used as the sample compressed content feature; Determine the product of the dimension-increased matrix and the compressed content feature of the sample as the restored content feature of the sample; Input the restored content features of the sample and the i-th task prompt into the i-th task model to obtain the sample i-th task processing result output by the i-th task model, wherein the i-th task model is used to process the i-th task, the i-th task prompt is used to describe the task requirements of the i-th task through text, and i is a positive integer; According to the difference between the sample's i-th task processing result and the i-th task result true value, the matrix parameters of the dimensionality reduction matrix and the dimensionality increase matrix are updated.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Encoding the sample content by the shared encoder to obtain sample content features; Input the sample content feature and the i-th task prompt into the i-th task model to obtain the sample i-th task processing result output by the i-th task model, wherein the i-th task model is used to process the i-th task, the i-th task prompt is used to describe the task requirements of the i-th task through text, and i is a positive integer; According to the difference between the i-th processing result of the sample and the true value of the i-th task result, the model parameters of the shared encoder and each of the i-th task models are updated.
8. The method according to claim 7, characterized in that The method further comprises: In the case where there are new tasks other than the i-th tasks, training a new task model based on the sample content, wherein the parameters of the shared encoder are fixed during the training of the new task model; Determining the model performance of the newly added task model through a test sample set; When the performance of the model meets the performance requirement, in response to a request for processing a new task for the first content, inputting the first content feature and the new task prompt read from the cache into the new task model to obtain a new task processing result output by the new task model; When the model performance does not meet the performance requirement, model parameters of the shared encoder and the newly added task model are updated based on the training sample set so that the model performance meets the performance requirement.
9. The method according to any one of claims 1 to 8, characterized in that: The shared encoder includes a first attention network, and encoding the content through the shared encoder and caching the content features obtained by the encoding include: Encode the content through the first attention network to obtain a first K matrix and a first V matrix; Using the first K matrix and the first V matrix as the content feature cache; The step of inputting the first content feature and the first task prompt into a first task model to obtain a first task processing result output by the first task model includes: The first K matrix, the first V matrix, and the first task prompt obtained by encoding the first content through the first attention network are input into the first task model to obtain the first task processing result.
10. The method according to claim 9, characterized in that The task model includes a second attention network and a task processing network, and the first K matrix, the first V matrix, and the first task prompt obtained by encoding the first content by the first attention network are input into the first task model to obtain the first task processing result, including: Encode the first task prompt through the second attention network to obtain a second K matrix, a second V matrix and a Q matrix; By using the second attention network, a K matrix obtained by concatenating the first K matrix and the second K matrix, a V matrix obtained by concatenating the first V matrix and the second V matrix, and the Q matrix are operated to obtain an attention value; The attention value is input into the task processing network to obtain the first task processing result.
11. The method according to any one of claims 1 to 10, characterized in that: The shared encoder is an encoding network in a large language model, and the task model is a task network in the large language model; The encoding of the content by the shared encoder and caching of the encoded content features include: Encoding the content through the encoding network in the large language model, and caching the content features obtained by encoding; The step of inputting the first content feature and the first task prompt into a first task model to obtain a first task processing result output by the first task model includes: The first content feature and the first task prompt are input into the task network in the large language model to obtain the first task processing result.
12. The method according to any one of claims 1 to 11, characterized in that: The content is text content, and the task model is a text content classification model. Different task models are used to process text content classification tasks in different vertical domains; The step of inputting the first content feature and the first task prompt into the first task model to obtain the first task processing result output by the first task model includes: The first text content feature and the text content classification prompt of the first vertical domain are input into the first text content classification model to obtain a first text content classification result.
13. A task processing device, characterized in that: The device comprises: A cache module, used to encode content through a shared encoder and cache content features obtained by encoding, wherein the shared encoder is an encoder shared by different task models, and different task models are used to process different tasks, and the shared encoder is obtained based on joint training of different tasks; a reading module, configured to read a first content feature corresponding to the first content from a cache in response to a first task processing request for the first content; A processing module is used to input the first content feature and the first task prompt into a first task model to obtain a first task processing result output by the first task model, wherein the first task model is used to process a first task, the first task prompt is used to describe the task requirements of the first task through text, and different tasks correspond to different task prompts and different task models.
14. A computer device, characterized in that: The computer device includes a processor and a memory; the memory stores at least one computer instruction, and the at least one computer instruction is used to be executed by the processor to implement the task processing method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer instruction, and the computer instruction is loaded and executed by a processor to implement the task processing method according to any one of claims 1 to 12.
16. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the task processing method as described in any one of claims 1 to 12.