Training method and device for abnormal target recognition model based on multi-task prompts
Through the abnormal target recognition model of multi-task prompts, the image encoder, text encoder and joint scheduler optimization training is used to solve the problems of high error rate and poor generalization in traditional methods, and more efficient abnormal target recognition is achieved.
Patent Information
- Application Number
- CN202411673640.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-11-21
AI Technical Summary
Traditional abnormal object recognition methods have problems with high error rate and poor generalization, mainly due to false associations in image features and insufficient number of abnormal samples.
An exception target recognition model based on multi-task cues is adopted, visual and text cues are generated through image encoder and text encoder, false associations are filtered using multi-layer perceptrons with masks, and task loss weights are allocated through joint schedulers to perform cross-task causal knowledge transfer and optimization training.
Effectively reduce the error rate, improve the generalization of abnormal target recognition, alleviate interference between tasks, and improve the effect of cross-task knowledge transfer.
Smart Images

Figure CN119516558B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a training method and device for an abnormal target recognition model based on multi-task prompts. Background Art
[0002] Abnormal target recognition technology aims to automatically detect and identify targets that do not conform to conventional or preset patterns from complex image or video backgrounds. It is of great significance in network governance, security monitoring, military reconnaissance, industrial inspection and other fields.
[0003] Traditional methods for identifying unusual objects often rely on manually set rules and thresholds, which are not only inefficient but also difficult to adapt to complex and changing environments. In recent years, with the rise of technologies such as deep learning, unusual object recognition has made significant progress. Deep learning algorithms can automatically learn feature representations from large amounts of data and are highly robust to complex backgrounds and lighting conditions, providing new ideas and methods for unusual object recognition.
[0004] However, traditional abnormal target recognition methods have two shortcomings: on the one hand, existing methods mainly rely on statistical correlations in image features and fail to effectively solve the problem of false correlations in data. False correlations in image data can easily mislead the model to learn non-causal features and produce incorrect prediction results; on the other hand, the number of abnormal samples in abnormal target recognition tasks is often less than that of normal samples. Traditional single-task models are not easy to learn sufficient abnormal knowledge, which easily reduces the generalization of abnormal recognition models.
[0005] In summary, traditional abnormal target recognition has the problems of high error rate and poor generalization. Summary of the Invention
[0006] In view of this, the present invention provides a training method for an abnormal target recognition model based on multi-task prompts, an abnormal target recognition method based on multi-task prompts, a device, a storage medium and a program product to solve the problems of high error rate and poor generalization of traditional abnormal target recognition.
[0007] In a first aspect, the present invention provides a training method for an abnormal target recognition model based on multi-task prompts, wherein the abnormal target recognition model includes: an image encoder, a text encoder, a multi-layer perceptron with a mask, and a joint scheduler; the text encoder and the image encoder are the text encoder and the image encoder of a pre-trained vision-language model; the method includes:
[0008] Obtaining task samples belonging to a target number of preset recognition tasks to obtain a set of task samples; wherein the task samples include training image samples and their corresponding text label samples;
[0009] An image encoder is used to generate visual cue features based on multi-task causal cues and training image samples; a text encoder is used to generate text cue features based on multi-task causal cues and text label samples; a masked multi-layer perceptron is used to perform feature filtering on the visual cue features to obtain counterfactual visual cue features; the multi-task causal cues include task cues for each task sample and shared cues for at least some task samples;
[0010] Contrastive learning is used to determine the task loss based on the similarity between counterfactual visual cue features and textual cue features, as well as the similarity between counterfactual visual cue features and preset recognition categories. A joint scheduler is used to assign weights to the task losses of each task sample. The weights and task losses are used to update the parameters of the multi-task causal cue and masked multi-layer perceptron until the preset training end conditions are reached.
[0011] During the training process of the abnormal target recognition model, the visual cue features and text cue features of the task samples output by the image encoder and text encoder are learned through task sharing and task-specific cues. Then, causal intervention is used to remove false correlations in the visual cue features to obtain counterfactual visual cue features. By comparative learning of counterfactual visual cue features and text cue features, causal knowledge across tasks is effectively transferred, which alleviates knowledge deficiencies, improves the generalization of abnormal target recognition, and reduces the error rate.
[0012] In an optional implementation, a joint scheduler is used to assign weights to task losses of various task samples; and the weights and task losses are used to update parameters of a multi-task causal prompt and a masked multilayer perceptron until a preset training end condition is reached, specifically including:
[0013] A joint scheduler is used to calculate the causal effect of task samples through counterfactual reasoning based on the visual cue features of task samples and counterfactual visual cue features.
[0014] A joint scheduler is used to calculate the gradient of the task loss of the task sample with respect to the shared hint;
[0015] A joint scheduler is used to calculate the causal affinity matrix of other task samples to the target task sample based on the gradients of the target task sample and other task samples; wherein the target task sample is each task sample in the set of task samples, and the other task samples are the task samples in the set of task samples except the target task sample;
[0016] A joint scheduler is used to determine a joint scheduling weight for each task sample based on a causal affinity matrix and a causal effect; the joint scheduler is an iteratively optimized joint scheduler using a joint scheduling training dataset, wherein the joint scheduling training dataset includes task samples belonging to the identified category;
[0017] The multi-task joint loss is determined according to the joint scheduling weight and task loss, and the multi-task joint loss is used to update the parameters of the multi-task causal prompt and the masked multi-layer perceptron until the preset training end condition is reached.
[0018] The causal effects of the selected features are measured through counterfactual reasoning, and the causal affinity between tasks is measured by calculating the gradient relationship between tasks. The joint scheduling weights are adaptively adjusted based on the causal effects and causal affinity between tasks, and the joint scheduling scheme is determined according to the joint scheduling weights to achieve better multi-task training. The obtained abnormal target recognition model can effectively perceive the relationship between tasks, alleviate the interference between tasks, and improve the cross-task knowledge transfer.
[0019] In an optional implementation, a method for obtaining the joint scheduling weight includes:
[0020] A learnable parameter set is constructed based on the weights. The learnable parameters in the learnable parameter set are reference ratio parameters for the causal affinity matrix and the causal effect when the joint scheduler allocates weights.
[0021] The acquired joint scheduling training dataset is used to verify the abnormal target recognition model, and the task loss and target loss of the joint scheduling training dataset are calculated to obtain the target loss.
[0022] The learnable parameter set is adjusted based on the target loss to obtain the joint scheduling weight.
[0023] In order to alleviate inter-task interference and improve cross-task knowledge transfer according to the impact of inter-task relationships on multi-task joint training, a learnable parameter set is constructed through weights. The joint scheduling training dataset is used to verify the abnormal target recognition model, and the target loss is constructed. The optimal task scheduling weights are solved according to the target loss for scheduling based on the joint scheduler.
[0024] In an optional implementation, a joint scheduler is used to assign weights to task losses of various task samples; and the weights and task losses are used to update parameters of a multi-task causal prompt and a masked multilayer perceptron until a preset training end condition is reached, specifically including:
[0025] Construct a two-level optimization problem based on target loss and multi-task joint loss;
[0026] Alternately solve the two-layer optimization problem until the preset end condition is met, and obtain the optimal set of learnable parameters and the parameters of the multi-task causal prompt and the masked multi-layer perceptron;
[0027] Among them, the two-level optimization problem includes:
[0028]
[0029] Where, L dev (θ * (φ)) represents the joint scheduling training dataset D dev where L(θ, φ) represents the multi-task joint loss, φ represents the set of learnable parameters, and θ represents the target parameters, which include the parameters of multi-task causal cues and multi-layer perceptrons with masks.
[0030] The optimization problem of multi-task joint training is converted into a two-level optimization problem for solution, avoiding a large amount of parameter calculations, improving computational efficiency, and accelerating model convergence.
[0031] In an optional embodiment, the step of solving the two-level optimization problem includes:
[0032] Fixed set of learnable parameters, solving target parameters based on underlying optimization problem;
[0033] The target parameters are brought into the upper-level optimization problem. Based on the Cauchy implicit function theorem, the chain method is used to calculate the gradient of the target loss with respect to the learnable parameter set, and the learnable parameter set is updated according to the gradient of the target loss with respect to the learnable parameter set.
[0034] In an optional implementation, the pre-trained vision-language model uses a pre-trained CLIP model.
[0035] In a second aspect, the present invention provides a method for identifying abnormal targets based on multi-task prompts, the method for identifying abnormal targets based on multi-task prompts comprising:
[0036] Obtain the image to be recognized;
[0037] The image to be identified is input into the trained multi-task prompt-based abnormal target recognition model to obtain an abnormal target prediction label; the trained multi-task prompt-based abnormal target recognition model is trained according to any of the above-mentioned training methods for the multi-task prompt-based abnormal target recognition model.
[0038] In a third aspect, the present invention provides a training device for an abnormal target recognition model based on multi-task prompts, wherein the abnormal target recognition model includes: an image encoder, a text encoder, a multi-layer perceptron with a mask, and a joint scheduler; the text encoder and the image encoder are the text encoder and the image encoder of a pre-trained vision-language model; the training device for the abnormal target recognition model based on multi-task prompts includes:
[0039] A sample acquisition module is used to acquire task samples belonging to a target number of preset recognition tasks to obtain a set of task samples; wherein the task samples include training image samples and their corresponding text label samples;
[0040] A feature extraction module is configured to generate visual cue features using an image encoder based on multi-task causal cues and training image samples; generate text cue features using a text encoder based on multi-task causal cues and text label samples; and perform feature filtering on the visual cue features using a masked multi-layer perceptron to obtain counterfactual visual cue features; the multi-task causal cues include task cues for each task sample and shared cues for at least some task samples;
[0041] The prediction and update module is used to use contrastive learning to determine the task loss based on the similarity between the counterfactual visual cue features and the textual cue features, as well as the similarity between the counterfactual visual cue features and the preset recognition categories; use a joint scheduler to assign weights to the task losses of each task sample; and use the weights and task losses to update the parameters of the multi-task causal cue and the masked multi-layer perceptron until the preset training end conditions are reached.
[0042] In a fourth aspect, the present invention provides an abnormal target identification device based on multi-task prompts, the abnormal target identification device based on multi-task prompts comprising:
[0043] An input acquisition module is used to acquire the image to be recognized;
[0044] The model prediction module is used to input the image to be identified into the trained multi-task prompt-based abnormal target recognition model to obtain an abnormal target prediction label; the trained multi-task prompt-based abnormal target recognition model is trained according to any of the above-mentioned training methods for the multi-task prompt-based abnormal target recognition model.
[0045] In a fifth aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to thereby execute the training method of the abnormal target recognition model based on multi-task prompts of the above-mentioned first aspect or any corresponding embodiment thereof or the abnormal target recognition method based on multi-task prompts of the second aspect.
[0046] In a sixth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the training method for an abnormal target recognition model based on multi-task prompts of the above-mentioned first aspect or any corresponding embodiment thereof, or the abnormal target recognition method based on multi-task prompts of the second aspect.
[0047] In the seventh aspect, the present invention provides a computer program product comprising computer instructions, the computer instructions being used to enable a computer to execute the training method for an abnormal target recognition model based on multi-task prompts of the above-mentioned first aspect or any corresponding embodiment thereof, or the abnormal target recognition method based on multi-task prompts of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 1 is a flow chart of a method for training an abnormal target recognition model based on multi-task prompts according to an embodiment of the present invention;
[0050] Figure 2 1 is a flow chart of an abnormal target recognition method based on multi-task prompts according to an embodiment of the present invention;
[0051] Figure 3 is a structural block diagram of a training device for an abnormal target recognition model based on multi-task prompts according to an embodiment of the present invention;
[0052] Figure 4 is a structural block diagram of an abnormal target recognition device based on multi-task prompts according to an embodiment of the present invention;
[0053] Figure 5 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0055] According to an embodiment of the present invention, a method for training an abnormal target recognition model based on multi-task prompts is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0056] In this embodiment, a training method for an abnormal target recognition model based on multi-task prompts is provided, which can be used in the above-mentioned computer systems, such as desktop computers, tablet computers, servers, etc. The abnormal target recognition model includes: an image encoder, a text encoder, a multi-layer perceptron with a mask, and a joint scheduler; the text encoder and the image encoder are the text encoder and image encoder of a pre-trained vision-language model. Figure 1 is a flow chart of a method for training an abnormal target recognition model based on multi-task prompts according to an embodiment of the present invention, such as Figure 1 As shown, the process includes the following steps:
[0057] Step S101 , obtaining task samples belonging to a target number of preset recognition tasks to obtain a set of task samples; wherein the task samples include training image samples and their corresponding text label samples.
[0058] Specifically, the set of task samples Task samples including N target recognition tasks Among them, each task sample Includes training image samples and text label samples corresponding to the training image samples.
[0059] In step S102, an image encoder is used to generate visual cue features based on multi-task causal cues and training image samples; a text encoder is used to generate text cue features based on multi-task causal cues and text label samples; a multi-layer perceptron with a mask is used to perform feature filtering on the visual cue features to obtain counterfactual visual cue features; the multi-task causal cues include task cues for each task sample and shared cues for at least some task samples.
[0060] The image encoder is the image encoder of the pre-trained vision-language model, and the text encoder is the text encoder of the pre-trained vision-language model.
[0061] Multi-task causal cues are a collection of task samples The proposed method consists of multimodal cues shared by at least some task samples (denoted as shared cues) and N task-specific multimodal cues (denoted as task cues). Multi-task causal cues aim to share knowledge between different tasks and learn unique task characteristics.
[0062] For a given task sample The training image samples included therein are input into the image encoder, and the task-shared shared cues and task-specific task cues are taken as additional inputs of the image encoder to obtain visual cue features.
[0063] Accordingly, the task sample The included text label samples are input into the text encoder, and the task-shared shared prompts and task-specific task prompts are taken as additional inputs of the text encoder to obtain text prompt features.
[0064] The embodiment of the present invention is based on an image encoder and a text encoder. On the basis of the input task samples, it learns task-sharing and task-specific knowledge through the additional input of multi-task causal cues, thereby mining multi-task features and obtaining corresponding visual cue features and text cue features.
[0065] The embodiment of the present invention further includes a multi-layer perceptron with a mask as a causal intervention network for filtering visual cue features. The embodiment of the present invention obtains visual cue features through causal intervention Counterfactual visual cue features
[0066] In one embodiment, the causal intervention network includes N visual cue features Same-sized, task-specific learnable masks The element value of the mask belongs to (0,1), and the visual cue features can be filtered through learnable parameters.
[0067] The embodiment of the present invention uses a multi-layer perceptron with a mask to perform causal intervention on the visual cue features obtained by the image encoder, mines multi-task causal features, thereby reducing false associations and learning causal visual features.
[0068] In step S103, contrastive learning is used to determine the task loss based on the similarity between the counterfactual visual cue feature and the textual cue feature, as well as the similarity between the counterfactual visual cue feature and the preset recognition category; a joint scheduler is used to assign weights to the task losses of each task sample; and the weights and task losses are used to update the parameters of the multi-task causal cue and the masked multi-layer perceptron until the preset training end condition is reached.
[0069] After obtaining the counterfactual visual cue features and text cue features, the similarity between the counterfactual visual cue features and the text cue features is further calculated, and the similarity between the counterfactual visual cue features and the preset recognition categories is also calculated. A contrastive target is used as the task loss, and the counterfactual visual cue features are associated with the text labels. The task loss is obtained based on the similarity between the counterfactual visual cue features and the text cue features and the similarity between the counterfactual visual cue features and all preset recognition categories. Among them, the contrastive target refers to comparing the similarity between paired image and text embeddings and the similarity between unpaired image and text embeddings, maximizing the similarity of the true paired image and text embeddings while minimizing the similarity of incorrect pairings. It should be emphasized that paired image and text embeddings refer to the counterfactual visual cue features and the text cue features extracted from the corresponding task samples; unpaired image and text embeddings refer to the counterfactual visual cue features and the preset recognition categories.
[0070] Afterwards, a joint scheduler is used to assign weights to the task losses of each task sample. The causal task-cue scheduler aims to complete the causal scheduling of cue features and tasks. On the one hand, taking into account the different effects of different cue features on tasks, for example, some cue features easily make the model output correct predictions, while some features with false correlations do the opposite. On the other hand, the impact of inter-task relationships on multi-task joint training is dynamic, and it is necessary to effectively perceive the inter-task relationships, alleviate inter-task interference, and improve cross-task knowledge transfer. In order to solve the above problems, the present invention proposes a causal task-cue joint scheduler, which adaptively adjusts the weights of task losses according to the joint causal influence of cue features and task relationships.
[0071] Based on the joint scheduler, on the basis of weights and task losses, learning is done to maximize the model's output of correct predictions, thereby updating the parameters of the multi-task causal hint and the masked multi-layer perceptron until the preset training end condition is reached.
[0072] It should be noted that the training end condition includes a specified number of training rounds or a model evaluation index threshold, which is not limited in the present invention.
[0073] The training method of the abnormal target recognition model based on multi-task prompts in the embodiment of the present invention sets the abnormal target recognition model to include an image encoder, a text encoder, a masked multi-layer perceptron and a joint scheduler, and uses the image encoder and the text encoder to learn task sharing and task-specific prompts to mine the visual prompt features and text prompt features of the task samples, and then removes the false correlation in the visual prompt features through causal intervention to obtain counterfactual visual prompt features. The counterfactual visual prompt features and text prompt features are used for comparative learning to effectively transfer causal knowledge across tasks, alleviate knowledge deficiencies, improve the generalization of abnormal target recognition, and reduce the error rate.
[0074] Furthermore, in an optional embodiment, the pre-trained vision-language model uses a pre-trained CLIP model. This embodiment further describes the above steps S101-S103 using the CLIP model as an example, but does not limit the present invention.
[0075] Specifically, for the set of task samples Initialize a task-shared multimodal prompt and N task-specific multimodal cues To share knowledge between different tasks and learn unique task characteristics. and Represent task-shared and task-specific cues for text, respectively; and They are task-shared and task-specific cues for images, respectively.
[0076] For a given task sample The task-shared and task-specific prompts are given as additional inputs to the image / text encoder, and the text encoder and image encoder Each consists of k Transform layers, where the text encoder and image encoder The k-th layer encoding process is shown in formula (1) and formula (2):
[0077]
[0078] Among them, [-, -, W k ] and [c k ,-,-,E k ] represents the connection operation of the k-th layer encoding process, - represents the hint of the middle layer, W k Represents the weight matrix of the kth Transform layer of the text encoder, E k represents the weight matrix of the kth Transform layer of the image encoder, c kRepresents the visual label representation of the k-th Transform layer of the image encoder; is the k-th layer encoding of the text encoder; is the k-th layer code of the image encoder; and Represent task-shared and task-specific cues for text, respectively; and They are task-shared and task-specific cues for images, respectively.
[0079] Finally, the CLIP image encoder is constructed by representing the visual label c in the last layer. K Input to the visual projection layer (ImageProj) to obtain visual cue features The process is shown in formula (3). CLIP text encoder represents the category token of the last Transform layer Input to the text projection layer (TextProj) to obtain the text label z. The process is shown in formula (4).
[0080]
[0081]
[0082] For tasks Obtaining counterfactual visual cue features through causal intervention with masked multi-layer perceptron The specific process is shown in formula (5):
[0083]
[0084] Among them, MLP is a multi-layer perceptron. is the N visual cue features included in the causal intervention network Same-sized, task-specific learnable masks After obtaining the counterfactual visual cue features, the task loss is constructed as shown in formula (6):
[0085]
[0086] in, is the predicted text label, which specifically refers to the text prompt feature output by the text encoder, z c is the visual label representation of the c-th category, C is the number of categories, sim(.) is the cosine similarity score, and τ is the temperature parameter of the clip model, is the counterfactual visual cue feature.
[0087] The embodiment of the present invention uses the open source CLIP model as a pre-trained multimodal large model, and transfers the pre-trained knowledge of the CLIP model to multiple anomaly recognition tasks by learning task-specific task prompts. At the same time, learning task-shared shared prompts realizes cross-task knowledge transfer to alleviate the problem of knowledge deficiency. By fine-tuning the prompts of the CLIP model to train the abnormal target recognition model to realize abnormal target recognition, the abnormal target recognition effect in multiple scenarios such as zero-sample and small-sample can be effectively improved.
[0088] In some optional implementations of this embodiment, a joint scheduler is used to assign weights to task losses for each task sample; the weights and task losses are used to update the parameters of the text encoder, image encoder, and masked multilayer perceptron until a preset training end condition is reached, specifically including the following steps:
[0089] In step S1031 , a joint scheduler is used to calculate the causal effect of the task sample through counterfactual reasoning based on the visual cue features and counterfactual visual cue features of the task sample.
[0090] Specifically, in order to quantify the impact of the selected cue features on the task, a causal effect submodule of the cue features is constructed in the scheduler to calculate the causal effect of the selected features through counterfactual reasoning.
[0091] More specifically, based on the above embodiment, given a task sample Visual cue features and counterfactual visual cue features after causal intervention The causal effect of the selected prompt feature is calculated as shown in formula (7):
[0092]
[0093] Where σ is the sigmoid function, yes Corresponding task samples mission loss, yes Corresponding task samples mission loss, when When the visual cue features In contrast, counterfactual visual cue features Can produce smaller task loss, which means that the selected features have a positive impact on correct prediction. When , the selected features have a negative impact on the task. As model training progresses, the value of the causal effect is expected to continue to increase (or at least not decrease) so that the model can learn more causal features.
[0094] In step S1032 , a joint scheduler is used to calculate the gradient of the task sample with respect to the shared hint based on the task loss.
[0095] Specifically, based on the task loss, mathematical knowledge is used to calculate the gradient of the task sample with respect to the shared hint.
[0096] More specifically, based on the above embodiment, given two task samples and Their loss functions are and Their gradients with respect to the shared cue are and
[0097] Step S1033, using a joint scheduler to calculate the causal affinity matrix of other task samples to the target task sample based on the gradients of the target task sample and other task samples; wherein the target task sample is each task sample in the set of task samples, and the other task samples are task samples in the set of task samples other than the target task sample.
[0098] Specifically, we construct an inter-task causal affinity submodule in the scheduler to measure inter-task causal affinity based on inter-task gradient relationships. Considering that the direction and magnitude of gradients can affect the joint training of multiple tasks, we quantify the inter-task causal affinity matrix A by calculating the relationship between their gradient directions and magnitudes.
[0099] More specifically, based on the above embodiment, A i,j Is the source task Target tasks The causal affinity of is shown in formula (8):
[0100]
[0101] in, and For two task samples, It is a sample task The loss function is It is a sample task The loss function is The gradient of the shared cue is The gradient of the shared cue is A i,j ∈[-1, 1], the numerator is the similarity of the gradient direction between tasks, and the denominator is the L2 norm of the gradient of the target task. i,j The larger the positive value of Task The greater the positive impact, the negative value indicates a negative impact. A is an asymmetric matrix with a main diagonal of 0. Task The causal affinity of A j,i , and A i,j have the same numerator, and the denominator is The L2 norm of the gradient of .
[0102] In step S1034, a joint scheduler is used to determine the joint scheduling weight of each task sample based on the causal affinity matrix and the causal effect; the joint scheduler is a joint scheduler optimized and iterated using a joint scheduling training data set, wherein the joint scheduling training data set includes task samples belonging to the identification category.
[0103] Specifically, a joint scheduling target submodule is constructed in the scheduler to schedule the learning weights of task samples. In order to jointly schedule prompts and tasks, the embodiment of the present invention is a task sample τ i Designed learnable weights This weight takes into account both the causal effect of the cue features and the causal affinity between tasks. The purpose of this is to assign larger learning weights to causal features (with greater causal effects) and positively correlated tasks (with greater causal affinity between tasks), while assigning smaller learning weights to non-causal features and negatively correlated tasks. The joint scheduling weights determined in step 1034 are the optimal scheduling weights determined by the joint scheduler after iterative optimization using the joint scheduling training dataset.
[0104] In more detail, based on the above embodiment, the joint scheduling weight As shown in formula (9):
[0105]
[0106] in, is the visual cue feature, b is the batch index of the joint scheduler to learn the joint scheduling weight, σ represents the sigmoid function, and is a learnable parameter, For task T in matrix A i A corresponding column of weight vectors, that is, each task to task A vector consisting of the affinities.
[0107] Step S1035: determine the multi-task joint loss according to the joint scheduling weight and the task loss, and use the multi-task joint loss to update the parameters of the multi-task causal prompt and the masked multi-layer perceptron until the preset training end condition is reached.
[0108] Finally, the multi-task joint loss can be determined based on the joint scheduling weight and task loss, and the joint scheduling target can be established based on the multi-task joint loss to perform multi-task learning, thereby updating the parameters of the text encoder, image encoder, and masked multi-layer perceptron.
[0109] More specifically, based on the above embodiment, based on the task scheduling weight The joint scheduling objective of multi-task learning (i.e., multi-task joint loss) is shown in formula (10):
[0110]
[0111] Among them, D train Representation task sample The training set is a training data set extracted from the set of task samples. It is a sample task The loss function for training at batch index b is, It is a sample task Counterfactual visual cue features.
[0112] Furthermore, in order to enable the joint scheduler to perform scheduling learning of task samples based on the joint scheduling weights, the joint scheduler is optimized and iterated using a joint scheduling training dataset, wherein the joint scheduling training dataset includes task samples belonging to the identification category.
[0113] Furthermore, the embodiments of the present invention do not impose any specific restrictions on the joint scheduling training dataset; for example, it may be a set of task samples extracted from a set of task samples that have not participated in training. The goal of the joint scheduler's optimization iterations is to find the optimal task scheduling weights. In some embodiments, the joint scheduler's optimization iterations can be performed simultaneously with model training, as further described in subsequent embodiments of the present invention.
[0114] The embodiment of the present invention measures the causal effect of the selected features through counterfactual reasoning, and measures the causal affinity between tasks by calculating the gradient relationship between tasks. The joint scheduling weight is adaptively adjusted based on the causal effect and the causal affinity between tasks, and the joint scheduling scheme is determined according to the joint scheduling weight to achieve better multi-task training. The obtained abnormal target recognition model can effectively perceive the relationship between tasks, achieve the effect of alleviating interference between tasks, and improve cross-task knowledge transfer.
[0115] Based on the above embodiment, in some optional implementations of this embodiment, a parameter solution module based on two-level optimization is constructed to synchronize the optimization iteration of the joint scheduler with the training of the model. On this basis, the method for obtaining the joint scheduling weight includes the following steps:
[0116] Step a1: Construct a learnable parameter set based on the weights; the learnable parameters in the learnable parameter set are reference ratio parameters for allocating the causal affinity matrix and the causal effect when the joint scheduler allocates weights.
[0117] Based on the above embodiments, it can be understood that the joint scheduler is intended to provide an appropriate joint scheduling solution to achieve better multi-task training. It is based on the causal affinity matrix, causal effects and learnable parameters and Calculated. On this basis, a set of learnable parameters can be constructed based on the weights Where T is the set of tasks. and Causal effect And the allocation reference ratio parameter corresponding to the causal affinity matrix A in the process of allocating weights.
[0118] Step a2: Use the acquired joint scheduling training dataset to verify the abnormal target recognition model, calculate the task loss and target loss of the joint scheduling training dataset.
[0119] Specifically, since the scheduling goal of this embodiment is to optimize the learnable parameter set To reduce the multi-task joint loss L. In order to optimize the parameter φ during the scheduling process, this embodiment introduces a small verification dataset (also known as the joint scheduling training dataset) Among them, B is all batches in the validation set, b is the current batch index, The verification samples and labels of the current batch b in the verification process are used in D dec The objective loss on φ is used to optimize the parameter φ in order to achieve the optimal scheduling weights for multi-task training.
[0120] In some embodiments, the joint scheduling training dataset is a validation set D from the model v A subset extracted from D dev The objective loss is used to optimize the parameter φ so that v The optimal scheduling weight for multi-task training is achieved on The validation set D v and task samples The training set D train You can use the task sample The task samples in are allocated according to the preset allocation ratio.
[0121] Thus, we can get the validation data set D dev The loss on , recorded as the target loss:
[0122] Step a3: Adjust the learnable parameter set based on the target loss to obtain the joint scheduling weight.
[0123] By adjusting the learnable parameter set with the goal of minimizing the target loss, the optimal scheduling weight for multi-task training, namely the joint scheduling weight, can be obtained.
[0124] In order to alleviate inter-task interference and improve cross-task knowledge transfer based on the impact of inter-task relationships on multi-task joint training, the embodiments of the present invention construct a learnable parameter set through weights, use a joint scheduling training dataset to verify the abnormal target recognition model, construct a target loss, and solve the optimal task scheduling weight based on the target loss for scheduling based on a joint scheduler.
[0125] Based on the above embodiment, in some optional implementations of this embodiment, a parameter solution module based on two-level optimization is constructed to synchronize the optimization iteration of the joint scheduler with the training of the model. On this basis, the optimization iteration of the joint scheduler and the training of the model specifically include the following steps:
[0126] Step b1: Construct a two-level optimization problem based on the target loss and the multi-task joint loss;
[0127] Step b2: Alternately solve the two-layer optimization problem until the preset end condition is met, and obtain the optimal set of learnable parameters and the parameters of the multi-task causal prompt and the masked multi-layer perceptron;
[0128] Specifically, based on the target loss and the multi-task joint loss, the multi-task joint optimization problem can be formulated as the following two-level optimization problem:
[0129]
[0130] Where φ represents the set of learnable parameters, θ represents the target parameters, and the target parameters include the parameters of multi-task causal cues and multi-layer perceptrons with masks; * Represents the updated φ; θ * Indicates the updated target parameter; st indicates that the subsequent formula is a constraint condition; Indicates that in the validation data set D dev The loss on . L(θ, φ) is the scheduling training loss (i.e., multi-task joint loss) defined in Equation (10), where φ is used for parameterization.
[0131] Solving the two-layer optimization problem until the preset termination condition is met yields the optimal set of learnable parameters and the parameters of the text encoder, image encoder, and masked multilayer perceptron. The optimal set of learnable parameters corresponds to the joint scheduling weights.
[0132] The embodiment of the present invention transforms the optimization problem of multi-task joint training into a two-layer optimization problem for solution, thereby avoiding a large amount of parameter calculation, improving computational efficiency, and accelerating model convergence.
[0133] In some optional implementations of this embodiment, two submodules, inner layer (lower layer) optimization and outer layer (upper layer) optimization, are constructed in the parameter solution module based on two-level optimization to solve the two-level optimization problem. Specifically, the steps of solving the two-level optimization problem include:
[0134] Step c1: Fix the set of learnable parameters and solve the target parameters based on the underlying optimization problem;
[0135] Specifically, the inner (lower) layer optimization submodule solves the inner (lower) layer optimization problem. In the actual solution process, a fixed parameter φ is used to update the parameter θ of the abnormal target recognition model. The parameter θ can be denoted as the target parameter, which includes the parameters of the aforementioned multi-task causal hint and the masked multilayer perceptron.
[0136] More specifically, the update of θ is performed by using the weighted gradient sum of samples in different tasks, as shown in formula (12):
[0137]
[0138] Among them, L(θ, φ) is the multi-task joint loss, φ represents the set of learnable parameters, θ represents the target parameter, It is a sample task The loss function for training on batch b is, It is a sample task The counterfactual visual cue features, is the joint scheduling weight.
[0139] Step c2: Bring the target parameters into the outer (upper) optimization problem, and based on the Cauchy implicit function theorem, use the chain method to calculate the gradient of the target loss with respect to the learnable parameter set, and update the learnable parameter set according to the gradient of the target loss with respect to the learnable parameter set.
[0140] Specifically, the outer (upper) optimization submodule is used to calculate the upper optimization problem. The outer (upper) optimization target needs to calculate the target loss L dev (θ * (φ)) with respect to the gradient of the hyperparameter φ. Since L dev (θ *The dependence of (φ) on φ is indirect and can be realized through the parameter θ, so the present invention uses implicit differentiation to obtain its implicit gradient.
[0141] In more detail, based on the Cauchy implicit function theorem, the loss function L can be derived using the chain rule dev (θ * The gradient of (φ) with respect to the hyperparameter φ is shown in formula (13):
[0142]
[0143] Among them, L is L(θ, φ), which represents the multi-task joint loss, φ represents the learnable parameter set, θ represents the target parameter, and L dev (θ * (φ) represents the validation data set D dev on the loss.
[0144] Furthermore, for deep neural network models, due to the huge parameter size and complex solution, it is usually computationally infeasible to directly calculate the inverse of the Hessian matrix in formula (13). To solve this problem, the present invention uses an H-times truncated Neumann sequence to approximate the inverse of the Hessian matrix, as shown in formula (14). By using this approximation method to approximate the inverse of the Hessian matrix, the implicit gradient can be effectively estimated. The process is shown in formula (15).
[0145]
[0146] Among them, I represents the identity matrix, L is L(θ, φ), which represents the multi-task joint loss, φ represents the learnable parameter set, θ represents the target parameter, and L dev Indicates that in the validation data set D dev on the loss.
[0147] Table 1 shows the accuracy comparison of different methods on dataset 1 under different proportions (10% and 20% training visible). It can be seen from Table 1 that the accuracy of the abnormal target recognition model obtained by the training method of the abnormal target recognition model based on multi-task prompts provided by the present invention is much better than the accuracy of the model trained by Zero (zero-shot learning, zero-sample learning method), CoOp (Context Optimization, a method for adapting visual-language models), VPT (Visual Prompt Tuning, a method for adjusting pre-trained models for visual tasks), MaPLe (Multi-modal Prompt Learning, a multimodal prompt learning method), Adapter (model fine-tuning), and MmAP (Multi-modal Alignment Prompt for Cross-domain Multi-task Learning) in the related art.
[0148] Table 1 Comparison of the accuracy of different methods on dataset 1 at different ratios (10% and 20% training visible)
[0149]
[0150] Table 2 shows the accuracy comparison of different methods on dataset 2 under a small sample setting (k-shot). According to the courseware in Table 2, the accuracy of the abnormal target recognition model obtained by the training method of the abnormal target recognition model based on multi-task prompts provided by the present invention under a small sample setting (k-shot) is far better than the accuracy of the model trained by COOp (Context Optimization, a method for adapting visual-language models), VPT (Visual Prompt Tuning, a method for adjusting pre-trained models for visual tasks), MaPLe (Multi-modal Prompt Learning, a multimodal prompt learning method), MLVPT (Multitask Vision-Language Prompt Tuning, a method for training large pre-trained visual-language models), Adapter (model fine-tuning), and MmAP (Multi-modal Alignment Prompt for Cross-domain Multi-task Learning) in the related art.
[0151] Table 2 Comparison of the accuracy of different methods on dataset 2 under small sample settings (k-shot)
[0152]
[0153] In this embodiment, a method for identifying abnormal targets based on multi-task prompts is provided, which can be used for the above-mentioned desktop computers, notebooks, servers, etc. Figure 2 is a flow chart of an abnormal target recognition method based on multi-task prompts according to an embodiment of the present invention, such as Figure 2 As shown, the process includes the following steps:
[0154] Step S201, obtaining an image to be recognized;
[0155] In step S202, the image to be identified is input into the trained abnormal target recognition model based on multi-task prompts to obtain an abnormal target prediction label; the trained abnormal target recognition model based on multi-task prompts is trained according to any of the training methods for the abnormal target recognition model based on multi-task prompts described above.
[0156] The abnormal target recognition method based on multi-task prompts in this embodiment is trained using the training method of the abnormal target recognition model based on multi-task prompts in the above embodiment, so that abnormal targets in the input image to be recognized can be effectively recognized.
[0157] In this embodiment, a training device for an abnormal target recognition model based on multi-task prompts is also provided. The device is used to implement the above-mentioned embodiments and preferred embodiments, and the details that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0158] This embodiment provides a training device for an abnormal target recognition model based on multi-task prompts, the abnormal target recognition model includes: an image encoder, a text encoder, a multi-layer perceptron with a mask, and a joint scheduler; the text encoder and the image encoder are the text encoder and the image encoder of the pre-trained vision-language model, such as Figure 3 As shown, the training device of the abnormal target recognition model based on multi-task prompts includes:
[0159] The sample acquisition module 301 is used to acquire task samples belonging to a target number of preset recognition tasks to obtain a set of task samples; wherein the task samples include training image samples and their corresponding text label samples;
[0160] Feature extraction module 302 is configured to generate visual cue features using an image encoder based on the multi-task causal cues and training image samples; generate text cue features using a text encoder based on the multi-task causal cues and text label samples; and perform feature filtering on the visual cue features using a masked multi-layer perceptron to obtain counterfactual visual cue features; the multi-task causal cues include task cues for each task sample and shared cues for at least some task samples;
[0161] The prediction and update module 303 is used to use contrastive learning to determine the task loss based on the similarity between the counterfactual visual cue features and the textual cue features, as well as the similarity between the counterfactual visual cue features and the preset recognition categories; use a joint scheduler to assign weights to the task losses of each task sample; and use the weights and task losses to update the parameters of the multi-task causal cue and the masked multi-layer perceptron until the preset training end conditions are reached.
[0162] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0163] The training device of the abnormal target recognition model based on multi-task prompts in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0164] This embodiment also provides an abnormal target recognition device based on multi-task prompts, such as Figure 4 As shown, the abnormal target recognition device based on multi-task prompts includes:
[0165] Input acquisition module 401, used to acquire the image to be recognized;
[0166] The model prediction module 402 is used to input the image to be identified into the trained multi-task prompt-based abnormal target recognition model to obtain an abnormal target prediction label; the trained multi-task prompt-based abnormal target recognition model is trained according to any of the above-mentioned training methods for the multi-task prompt-based abnormal target recognition model.
[0167] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0168] The abnormal target identification device based on multi-task prompts in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0169] The embodiment of the present invention also provides a computer device having the above Figure 3 or Figure 4 The training device of the abnormal target recognition model based on multi-task prompts or the abnormal target recognition device based on multi-task prompts shown.
[0170] See also Figure 5 , Figure 5 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.
[0171] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0172] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0173] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0174] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0175] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means.
[0176] The input device 30 can receive input digital or character information and generate key signal input related to user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0177] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0178] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0179] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A training method for an abnormal target recognition model based on multi-task prompts, characterized in that: The abnormal target recognition model includes: an image encoder, a text encoder, a multi-layer perceptron with a mask, and a joint scheduler; the text encoder and the image encoder are the text encoder and image encoder of a pre-trained vision-language model; the method includes: Acquire task samples belonging to a target number of preset recognition tasks to obtain a set of task samples; wherein the task samples include training image samples and their corresponding text label samples; The image encoder is used to generate visual cue features based on the multi-task causal cues and the training image samples; the text encoder is used to generate text cue features based on the multi-task causal cues and the text label samples; the masked multi-layer perceptron is used to perform feature filtering on the visual cue features to obtain counterfactual visual cue features; the multi-task causal cues include task cues for each of the task samples and shared cues for at least some of the task samples; Contrastive learning is used to determine the task loss based on the similarity between the counterfactual visual cue feature and the textual cue feature, as well as the similarity between the counterfactual visual cue feature and the preset recognition category; the joint scheduler is used to assign weights to the task losses of each of the task samples; and the weights and the task losses are used to update the parameters of the multi-task causal cue and the masked multi-layer perceptron until the preset training end condition is reached.
2. The method according to claim 1, characterized in that The step of using the joint scheduler to assign weights to task losses of the task samples; and using the weights and the task losses to update the multi-task causal prompt and the parameters of the masked multi-layer perceptron until a preset training end condition is reached, specifically includes: The joint scheduler is used to calculate the causal effect of the task sample through counterfactual reasoning based on the visual cue feature and the counterfactual visual cue feature of the task sample; Calculating a gradient of the task loss of the task sample with respect to the shared hint using the joint scheduler; The joint scheduler is used to calculate the causal affinity matrix of the other task samples to the target task sample based on the gradients of the target task sample and the other task samples; wherein the target task sample is each task sample in the set of task samples, and the other task samples are task samples in the set of task samples other than the target task sample; Determining a joint scheduling weight for each of the task samples using the joint scheduler according to the causal affinity matrix and the causal effect; wherein the joint scheduler is an iteratively optimized joint scheduler using a joint scheduling training dataset, wherein the joint scheduling training dataset includes task samples belonging to the identified category; A multi-task joint loss is determined according to the joint scheduling weight and the task loss, and the multi-task joint loss is used to update the multi-task causal prompt and the parameters of the multi-layer perceptron with mask until a preset training end condition is reached.
3. The method according to claim 2, characterized in that The joint scheduling weight is obtained by: Constructing a learnable parameter set according to the weights; the learnable parameters in the learnable parameter set are reference ratio parameters for allocating the causal affinity matrix and the causal effect when the joint scheduler allocates the weights; The abnormal target recognition model is verified using the acquired joint scheduling training data set, and the task loss and target loss of the joint scheduling training data set are calculated to obtain the target loss; The learnable parameter set is adjusted based on the target loss to obtain the joint scheduling weight.
4. The method according to claim 3, characterized in that The step of using the joint scheduler to assign weights to task losses of the task samples; and using the weights and the task losses to update the multi-task causal prompt and the parameters of the masked multi-layer perceptron until a preset training end condition is reached, specifically includes: Constructing a two-level optimization problem based on the target loss and the multi-task joint loss; Alternately solving the two-layer optimization problem until a preset end condition is met, thereby obtaining an optimal set of learnable parameters and parameters of the multi-task causal prompt and the masked multi-layer perceptron; The two-layer optimization problem includes: Where, L dev (θ*(φ)) represents the joint scheduling training dataset D dev where L(θ, φ) represents the multi-task joint loss, φ represents the set of learnable parameters, and θ represents the target parameters, which include the parameters of multi-task causal cues and multi-layer perceptrons with masks.
5. The method according to claim 4, characterized in that The steps of solving the two-level optimization problem include: Fixing the learnable parameter set and solving the target parameter based on the underlying optimization problem; The target parameters are brought into the upper-level optimization problem, and based on the Cauchy implicit function theorem, the chain method is used to calculate the gradient of the target loss with respect to the learnable parameter set, and the learnable parameter set is updated according to the gradient of the target loss with respect to the learnable parameter set.
6. The method according to claim 4, characterized in that The pre-trained vision-language model adopts a pre-trained CLIP model.
7. A method for identifying abnormal targets based on multi-task prompts, characterized in that: The method comprises: Obtain the image to be recognized; The image to be identified is input into the trained abnormal target recognition model based on multi-task prompts to obtain an abnormal target prediction label; the trained abnormal target recognition model based on multi-task prompts is trained according to the training method of the abnormal target recognition model based on multi-task prompts described in any one of claims 1-6.
8. A training device for an abnormal target recognition model based on multi-task prompts, characterized in that: The abnormal target recognition model includes: an image encoder, a text encoder, a multi-layer perceptron with a mask, and a joint scheduler; the text encoder and the image encoder are the text encoder and image encoder of a pre-trained vision-language model; the device includes: A sample acquisition module is used to acquire task samples belonging to a target number of preset recognition tasks to obtain a set of task samples; wherein the task samples include training image samples and their corresponding text label samples; a feature extraction module configured to generate visual cue features using the image encoder based on the multi-task causal cues and the training image samples; generate text cue features using the text encoder based on the multi-task causal cues and the text label samples; and perform feature filtering on the visual cue features using the masked multi-layer perceptron to obtain counterfactual visual cue features; the multi-task causal cues include task cues for each task sample and shared cues for at least some of the task samples; A prediction and update module is used to use contrastive learning to determine the task loss based on the similarity between the counterfactual visual cue feature and the textual cue feature, as well as the similarity between the counterfactual visual cue feature and the preset recognition category; use the joint scheduler to assign weights to the task losses of each of the task samples; and use the weights and the task losses to update the parameters of the multi-task causal cue and the masked multi-layer perceptron until a preset training end condition is reached.
9. An abnormal target recognition device based on multi-task prompts, characterized in that: The device comprises: An input acquisition module is used to acquire the image to be recognized; A model prediction module is used to input the image to be identified into a trained abnormal target recognition model based on multi-task prompts to obtain an abnormal target prediction label; the trained abnormal target recognition model based on multi-task prompts is trained according to the training method of the abnormal target recognition model based on multi-task prompts described in any one of claims 1-6.
10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the training method of the abnormal target recognition model based on multi-task prompts according to any one of claims 1 to 6, or executes the abnormal target recognition method based on multi-task prompts according to claim 7 by executing the computer instructions.
Citation Information
Patent Citations
Small sample visual classification method and device based on retrieval enhancement mechanism and visual cue learning
CN117953282A
Zero sample anomaly detection method based on multi-mode learnable prompt
CN118865000A