Training Method, Device, Electronic Device and Storage Medium of Resource Classification Model
By configuring sample resources and classification task identification for resource classification tasks, combining multi-head self-attention mechanism and multi-layer deep neural networks, the resource classification model is optimized, and the complex model structure is solved and more efficient classification prediction is achieved.
Patent Information
- Application Number
- CN202210023291.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-01-10
AI Technical Summary
The structure of the resource classification model in the prior art is too complex to effectively distinguish different classification tasks.
By configuring sample resources and classification task identifiers for each resource classification task and inputting them into the neural network model, the multi-head self-attention mechanism and multi-layer deep neural network of content feature vectors and conditional feature vectors are used to optimize the model structure and avoid configuring different output networks for different tasks.
The model structure is simplified, the prediction effect and scalability of the model are improved, and different classification tasks can be more accurately distinguished.
Smart Images

Figure CN114492601B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, an electronic device, and a storage medium for training a resource classification model. Background Art
[0002] In natural language processing tasks, there are often situations where multiple classification tasks are learned in parallel, that is, multiple classification learning tasks are integrated into a neural network model, and the neural network model is trained into a resource classification model. In related technologies, when training a resource classification model, different output networks are usually configured for different tasks to achieve the distinction of different tasks, which results in a very complex structure of the finally obtained resource classification model. Summary of the Invention
[0003] The present disclosure provides a method, an apparatus, an electronic device, and a storage medium for training a resource classification model, so as to at least solve the problem of the complex structure of the resource classification model in related technologies. The technical solutions of the present disclosure are as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, a method for training a resource classification model is provided, including: obtaining training samples of each resource classification task in multiple resource classification tasks; the training samples include sample resources of the corresponding resource classification task, classification task identifiers of the corresponding resource classification task, and label information, where the label information is used to indicate the reference category of the sample resources in the corresponding training samples; inputting the training samples of each resource classification task in the multiple resource classification tasks into a neural network model to obtain classification prediction results corresponding to the training samples; the classification prediction results are determined according to the content features of the sample resources in the training samples and the classification rules of the task types corresponding to the classification task identifiers; updating the parameters of the neural network model according to the classification prediction results corresponding to the training samples and the label information in the training samples; performing iterative training on the updated neural network model until the neural network model meets the model convergence condition, and determining the converged neural network model as the first resource classification model.
[0005] In a possible implementation, the training samples of each resource classification task among multiple resource classification tasks are input into a neural network model to obtain classification prediction results corresponding to the training samples, including: inputting the training samples of each resource classification task into the neural network model, and the neural network model performs the following steps: obtaining a content feature vector according to the sample resources of each resource classification task, where the content feature vector is used to characterize the content features of the sample resources of each resource classification task; obtaining a conditional feature vector according to the classification task identifier of each resource classification task, where the conditional feature vector is used to characterize the classification rules of the task type corresponding to the classification task identifier of each resource classification task; the dimension of the conditional feature vector is the same as that of the content feature vector; determining the classification prediction result corresponding to the training sample according to the content feature vector and the conditional feature vector.
[0006] In another possible implementation, the sample resources include multiple sample images, and the neural network model includes at least an image classification network and a self-attention network. Obtaining a content feature vector according to the sample resources of each resource classification task includes: inputting the multiple sample images into the image classification network for feature extraction to obtain the feature vector of each sample image; inputting the feature vector of each sample image into the self-attention network for feature interaction between every two sample images to obtain the content feature vector.
[0007] In another possible implementation, obtaining a conditional feature vector according to the classification task identifier of each resource classification task includes: determining the target number of rows according to the classification task identifier; determining that the content corresponding to the target number of rows in the pre-set dictionary matrix is the conditional feature vector.
[0008] In another possible implementation, the neural network model includes a multi-layer deep neural network. Determining the classification prediction result corresponding to the training sample according to the content feature vector and the conditional feature vector includes: performing a multi-head self-attention mechanism on the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; inputting the joint vector into the multi-layer deep neural network to obtain the classification prediction result corresponding to the training sample.
[0009] In another possible implementation, the neural network model includes a multi-layer deep neural network. Determining the classification prediction result corresponding to the training sample according to the content feature vector and the conditional feature vector includes: concatenating the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; inputting the joint vector into the multi-layer deep neural network to obtain the classification prediction result corresponding to the training sample.
[0010] In another possible implementation, the method further includes: performing multi-task joint training on the neural network model according to the training samples of each resource classification task and the training samples of the new resource classification task to obtain a second resource classification model; wherein, the training samples of the new resource classification task include the sample resources of the new resource classification task, the classification task identifier of the new resource classification task, and the new label information, and the second resource classification model is used to execute multiple resource classification tasks and the new resource classification task.
[0011] In another possible implementation, the method further includes: obtaining a prediction sample, where the prediction sample includes a prediction sample image and a prediction task identifier; inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction result.
[0012] In another possible implementation, inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction result includes: inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction probability; in the case where the prediction probability is greater than or equal to a preset threshold, determining that the prediction sample is a positive sample of the task type corresponding to the prediction task identifier.
[0013] In another possible implementation, obtaining the training samples of each resource classification task among multiple resource classification tasks includes: obtaining the training sample video of each resource classification task and the label information corresponding to the training sample video, and determining the classification task identifier of each resource classification task; performing the following steps for the training sample video of each resource classification task: obtaining the global image of the first preset number of frames of the training sample video; performing object detection on each frame of the global image to obtain the local image of the second preset number of frames, where the local image includes the object part area, and the global image of the first preset number of frames and the local image of the second preset number of frames are the sample resources of each resource classification task.
[0014] In another possible implementation, the method further includes: determining the number of rows of the matrix according to the number of tasks corresponding to multiple resource classification tasks; determining the number of columns of the matrix according to the dimension of the preset content feature vector; obtaining a dictionary matrix according to the number of rows of the matrix, the number of columns of the matrix, and the preset model.
[0015] In another possible implementation, the method further includes: during each round of iterative training of the neural network model, updating the dictionary matrix according to the classification prediction result corresponding to the training sample and the label information in the training sample; the updated dictionary matrix is used to determine the conditional feature vector during the next round of iterative training of the neural network model.
[0016] According to a second aspect of the embodiments of the present disclosure, there is provided a training device for a resource classification model, including: an acquisition module configured to acquire training samples of each resource classification task in a plurality of resource classification tasks; the training samples include sample resources of the corresponding resource classification task, a classification task identifier of the corresponding resource classification task, and label information, and the label information is used to indicate the reference category of the sample resources in the corresponding training samples; a training module configured to input the training samples of each resource classification task in the plurality of resource classification tasks into a neural network model to obtain a classification prediction result corresponding to the training samples; the classification prediction result is determined according to the content features of the sample resources in the training samples and the classification rules of the task types corresponding to the classification task identifiers; an update module configured to update the parameters of the neural network model according to the classification prediction result corresponding to the training samples and the label information in the training samples; an iteration module configured to perform iterative training on the updated neural network model until the neural network model meets the model convergence condition, and determine the converged neural network model as the first resource classification model.
[0017] In a possible implementation manner, the training module is specifically configured to perform: input the training samples of each resource classification task into the neural network model, and the neural network model performs the following steps: according to the sample resources of each resource classification task, obtain a content feature vector, and the content feature vector is used to characterize the content features of the sample resources of each resource classification task; according to the classification task identifier of each resource classification task, obtain a conditional feature vector, and the conditional feature vector is used to characterize the classification rules of the task types corresponding to the classification task identifiers of each sample resource task; the dimension of the conditional feature vector is the same as the dimension of the content feature vector; according to the content feature vector and the conditional feature vector, determine the classification prediction result corresponding to the training samples.
[0018] In another possible implementation manner, the sample resources include a plurality of sample images, and the neural network model includes at least an image classification network and a self-attention network. The training module is specifically configured to perform: input the plurality of sample images into the image classification network for feature extraction to obtain a feature vector of each sample image; input the feature vector of each sample image into the self-attention network for feature interaction between every two sample images to obtain a content feature vector.
[0019] In another possible implementation manner, the training module is specifically configured to perform: according to the classification task identifier, determine the target number of rows; determine that the content corresponding to the target number of rows in the preset dictionary matrix is the conditional feature vector.
[0020] In another possible implementation, the neural network model includes a multi-layer deep neural network, and the training module is specifically configured to perform: performing a multi-head self-attention mechanism on the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; inputting the joint vector into the multi-layer deep neural network to obtain a classification prediction result corresponding to the training sample.
[0021] In another possible implementation, the neural network model includes a multi-layer deep neural network, and the training module is specifically configured to perform: concatenating the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; inputting the joint vector into the multi-layer deep neural network to obtain a classification prediction result corresponding to the training sample.
[0022] In another possible implementation, the training module is further configured to perform: performing multi-task joint training on the neural network model according to the training samples of each resource classification task and the training samples of the new resource classification task to obtain a second resource classification model; wherein, the training samples of the new resource classification task include the sample resources of the new resource classification task, the classification task identifier of the new resource classification task, and the new label information, and the second resource classification model is used to perform multiple resource classification tasks and the new resource classification task.
[0023] In another possible implementation, the apparatus further includes a prediction module, configured to perform: obtaining a prediction sample, where the prediction sample includes a prediction sample image and a prediction task identifier; inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction result.
[0024] In another possible implementation, the prediction module is specifically configured to perform: inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction probability; in the case where the prediction probability is greater than or equal to a preset threshold, determining that the prediction sample is a positive sample of the task type corresponding to the prediction task identifier.
[0025] In another possible implementation, the acquisition module is specifically configured to perform: acquiring the training sample video of each resource classification task and the label information corresponding to the training sample video, and determining the classification task identifier of each resource classification task; performing the following steps for the training sample video of each resource classification task: acquiring the global image of the first preset number of frames of the training sample video; performing object detection on each frame of the global image to acquire the local image of the second preset number of frames, where the local image includes the object part area, and the global image of the first preset number of frames and the local image of the second preset number of frames are the sample resources of each resource classification task.
[0026] In another possible implementation, the apparatus further includes a configuration module configured to perform: determining the number of rows of the matrix according to the number of tasks corresponding to multiple resource classification tasks; determining the number of columns of the matrix according to the dimension of the preset content feature vector; and obtaining a dictionary matrix according to the number of rows of the matrix, the number of columns of the matrix, and a preset model.
[0027] In another possible implementation, the configuration module is further configured to perform: in each round of iterative training of the neural network model, updating the dictionary matrix according to the classification prediction result corresponding to the training sample and the label information in the training sample; and the updated dictionary matrix is used to determine the conditional feature vector during the next round of iterative training of the neural network model.
[0028] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein, the processor is configured to execute the instructions to implement the method according to the first aspect and any one of its possible implementations described above.
[0029] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by the processor of the electronic device, enabling the electronic device to execute the method according to the first aspect and any one of its possible implementations described above.
[0030] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, the computer program product includes computer instructions, when the computer instructions run on the electronic device, enabling the electronic device to execute the method according to the first aspect and any one of its possible implementations described above.
[0031] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: By configuring sample resources and classification task identifiers for the training samples of each resource classification task and inputting them into the neural network at the same time, it is possible to distinguish the classification tasks of different sample resources through different classification task identifiers, so that it is no longer necessary to configure different output networks for different classification tasks, solving the technical problem of complex model structure in the prior art. In addition, by jointly optimizing the neural network model through multiple classification prediction results of multiple resource classification tasks, the prediction effect of the model can also be improved.
[0032] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0034] Figure 1It is a flowchart of a method for training a resource classification model shown according to an exemplary embodiment;
[0035] Figure 2 It is a flowchart of another method for training a resource classification model shown according to an exemplary embodiment;
[0036] Figure 3 It is a flowchart of another method for training a resource classification model shown according to an exemplary embodiment;
[0037] Figure 4 It is a framework diagram of a method for training a resource classification model shown according to an exemplary embodiment;
[0038] Figure 5 It is a block diagram of a device for training a resource classification model shown according to an exemplary embodiment;
[0039] Figure 6 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation
[0040] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0042] Before introducing the method for training the resource classification model provided by the present disclosure in detail, the application scenarios and implementation environments involved in the present disclosure will be briefly introduced.
[0043] First, the application scenarios involved in the present disclosure will be briefly introduced.
[0044] In the related art, for different tasks or the same task in different regions, it is necessary to train as multiple resource classification tasks. For example, a model is trained separately for each resource classification task, or multiple resource classification tasks are directly integrated into a neural network model. Therefore, when training a resource classification model for multiple resource classification tasks, different output networks are usually configured for different tasks to achieve the distinction of different tasks, which results in a very complex structure of the finally obtained resource classification model.
[0045] Specifically, when training a resource classification model, shared mechanism training is mainly adopted, such as shared mechanisms like the hard sharing mode of parameters, the soft sharing mode of parameters, the hierarchical sharing mode, and the shared-private mode.
[0046] In view of the above problems, the present disclosure provides a training method for a resource classification model. By configuring sample resources and classification task identifiers for the training samples of each resource classification task and inputting them into a neural network simultaneously, it is possible to distinguish the classification tasks of different sample resources through different classification task identifiers, thus eliminating the need to configure different output networks for different classification tasks and solving the technical problem of the complex model structure in the prior art. In addition, by jointly optimizing the neural network model through multiple classification prediction results of multiple resource classification tasks, the prediction effect of the model can also be improved.
[0047] Secondly, the implementation environment (implementation architecture) related to the present disclosure will be briefly introduced below.
[0048] The training method for the resource classification model provided by the embodiments of the present disclosure can be applied to an electronic device. The electronic device can be a terminal device or a server. Among them, the terminal device can be a smart phone, a tablet computer, a handheld computer, a vehicle-mounted terminal, a desktop computer, a laptop computer, etc. The server can be any server or server cluster, and the present disclosure does not make any limitations in this regard.
[0049] In addition, it should be noted that the training sample information involved in the present disclosure (including but not limited to training sample video information, training sample image information, classification task identifiers, etc.) are all information authorized by the user or fully authorized by all parties.
[0050] For the sake of easy understanding, the task processing method provided by the present disclosure will be specifically introduced below with reference to the accompanying drawings.
[0051] Figure 1 is a flowchart of a training method for a resource classification model shown according to an exemplary embodiment, for an electronic device. As Figure 1 shown, the training method for the resource classification model includes the following steps:
[0052] In S101, training samples for each resource classification task among multiple resource classification tasks are obtained; the training samples include sample resources for the corresponding resource classification task, classification task identifiers for the corresponding resource classification task, and label information.
[0053] Among them, the label information is used to indicate the reference category of the training samples for each resource classification task.
[0054] Optionally, the reference category includes positive samples or negative samples.
[0055] In one implementation, the categories of the training samples for each resource classification task include positive samples and negative samples. That is, each resource classification task includes training samples of the positive sample category and training samples of the negative sample category. Among them, the training samples of the positive sample category include sample resources for the corresponding resource classification task, classification task identifiers for the corresponding resource classification task, and label information, and this label information is used to indicate that the training samples for the corresponding resource classification task are positive samples. The training samples of the negative sample category include sample resources for the corresponding resource classification task, classification task identifiers for the corresponding resource classification task, and label information, and this label information is used to indicate that the training samples for the corresponding resource classification task are negative samples.
[0056] In one implementation, the multiple resource classification tasks include tasks of multiple different task types. For example, the task of identifying whether there is a first content in the sample resource (referred to as the first task), the task of identifying whether there is a second content in the sample resource (referred to as the second task), and the task of identifying whether there is a third content in the sample resource (referred to as the third task) belong to tasks of different task types. Among them, the first content, the second content, and the third content are different contents.
[0057] In another implementation, the multiple resource classification tasks include tasks for different geographical regions, that is, a resource classification task is established for each different geographical region. Among them, the tasks for different geographical regions include tasks of the same task type for different geographical regions.
[0058] Optionally, different resource classification tasks have different training samples.
[0059] In S102, the training samples for each resource classification task among the multiple resource classification tasks are input into a neural network model to obtain classification prediction results corresponding to the training samples.
[0060] Among them, the classification prediction results are determined according to the content features of the sample resources in the training samples and the classification rules of the task types corresponding to the classification task identifiers.
[0061] Optionally, different task types have different classification rules. For example, the classification rule for the first task is to identify whether there is a first content on the sample resource, the classification rule for the second task is to identify whether there is a second content on the sample resource, and the classification rule for the third task is to identify whether there is a third content on the sample resource.
[0062] Optionally, the electronic device inputs the training samples of each resource classification task in the multiple resource classification tasks into the neural network model. For example, the multiple resource classification tasks include a first task, a second task, and a third task, that is, the training samples of the first task, the training samples of the second task, and the training samples of the third task are input into the neural network model.
[0063] In one implementation, taking the first task as an example for illustration, the first task includes a first training sample and a second training sample. Among them, the category of the first training sample is a positive sample, and the category of the second training sample is a negative sample. The neural network model determines the classification prediction result corresponding to the first training sample of the first task according to the sample resource of the first task and the classification task identifier of the first task in the first training sample, and determines the classification prediction result corresponding to the second training sample of the first task according to the sample resource of the first task and the classification task identifier of the first task in the second training sample. In addition, the label information in the first training sample is used to indicate that the first training sample is a positive sample, and the label information in the second training sample is used to indicate that the second training sample is a negative sample.
[0064] In one implementation, the training sample of each resource classification task is configured with the sample resource of each resource classification task and the classification task identifier of each resource classification task, and the sample resource of each resource classification task is associated with the classification task identifier of each resource classification task. That is, the classification task identifier of each resource classification task is used to represent the resource classification task to be performed on the sample resource of each resource classification task.
[0065] In one implementation, when inputting the training samples of each resource classification task into the neural network model, the sample resource and the classification task identifier of each resource classification task can be combined into a pair and input into the neural network model.
[0066] In S103, according to the classification prediction result corresponding to the training sample and the label information in the training sample, the parameters of the neural network model are updated.
[0067] In one implementation, after the electronic device obtains the classification prediction results corresponding to the training samples of each resource classification task, it determines the loss function values of the training samples of each resource classification task according to the classification prediction results corresponding to the training samples of each resource classification task and the label information in the training samples, and further obtains the loss function values of multiple resource classification tasks. Further, the electronic device determines the average loss function value of the multiple loss function values and updates the parameters of the neural network model.
[0068] In another implementation, the electronic device calculates the loss function value based on the classification prediction result corresponding to the training sample of each resource classification task and the label information of the training sample of each resource classification task. For example, the cross-entropy loss function is used to determine the loss function value. The label information in the training sample includes an identifier for characterizing the reference category of the training sample. Taking the first task as an example, the label information of the training sample of the first task can be the sample identifier "1", indicating that the training sample is a positive sample, that is, the training sample includes the first content, and the sample identifier "0" indicates that the training sample is a negative sample, that is, the training sample does not include the first content. The classification prediction result corresponding to the training sample can be "yes" or "no". Taking the first task as an example, when the classification prediction result is "yes", it indicates that the sample resource of the training sample includes the first content, and when the classification prediction result is "no", it indicates that the sample resource of the training sample does not include the first content.
[0069] Optionally, when updating the parameters of the neural network model based on the classification prediction result corresponding to the training sample and the label information in the training sample, the method of backpropagation and / or gradient descent can be used.
[0070] In S104, iterative training is performed on the updated neural network model until the neural network model meets the model convergence condition, and the converged neural network model is determined as the first resource classification model.
[0071] In one implementation, during each round of iterative training of the neural network model, the training samples of each resource classification task are the same. At this time, the above S102 - S103 are iteratively executed on the updated neural network model to implement iterative training of the neural network model.
[0072] In another implementation, during each round of iteration of the neural network model, the training samples of each resource classification task are different. At this time, the above S101 - S103 are iteratively executed on the updated neural network model to implement iterative training of the neural network model.
[0073] In one implementation, the model convergence condition includes the convergence of the loss function value of the classification prediction result, that is, the difference between the loss function values of the classification preset results of two adjacent times is less than or equal to a preset threshold.
[0074] In another implementation, the model convergence condition includes that the number of times of iteratively executing the above steps on the neural network model is greater than or equal to a preset number of iterations.
[0075] It should be noted that the present application realizes the multi-task joint training of the neural network model through the above S101-S104, so as to obtain a first resource classification model capable of executing multiple resource classification tasks.
[0076] In the above embodiments, by configuring sample resources and classification task identifiers for the training samples of each resource classification task and inputting them into the neural network at the same time, it is possible to distinguish the classification tasks of different sample resources through different classification task identifiers, so that there is no need to configure different output networks for different classification tasks, solving the technical problem of complex model structure in the prior art. In addition, by jointly optimizing the neural network model through the classification prediction results of multiple resource classification tasks, the prediction effect of the model can also be improved.
[0077] In a possible implementation, in combination with Figure 1 , as Figure 2 shown, S102 includes:
[0078] S102a: Input the training samples of each resource classification task into the neural network model, and execute the following S102b-S102d through the neural network model.
[0079] S102b: Obtain a content feature vector according to the sample resources of each resource classification task, and the content feature vector is used to characterize the content features of the sample resources of each resource classification task.
[0080] Optionally, the neural network model extracts features from the sample resources of each resource classification task to obtain the content feature vectors of the sample resources of each resource classification task.
[0081] In one implementation, the electronic device pre-determines that the dimension of the content feature vector is D, and D is a positive integer. For example, the sample resources include M sample images (M is a positive integer), and the neural network model extracts features from the M sample images respectively to obtain M feature vectors, and the dimension of each feature vector is D. Based on this, an M*D-dimensional content feature vector of the sample resources of each resource classification task is obtained.
[0082] S102c: Obtain a conditional feature vector according to the classification task identifier of each resource classification task, and the conditional feature vector is used to characterize the classification rules of the task type corresponding to the classification task identifier of each resource classification task.
[0083] Among them, the dimension of the conditional feature vector is the same as that of the content feature vector. By configuring the conditional feature vector and the content feature vector with the same dimension, when correlating the conditional feature vector and the content feature vector, the correlation between the conditional feature vector and the content feature vector can be improved, thereby improving the accuracy of determining the classification prediction result according to the content feature vector and the conditional feature vector.
[0084] Optionally, according to the target number of rows determined by the classification task identifier, an Embedding vector at the position corresponding to the target number of rows is obtained from a pre-set dictionary matrix to obtain a conditional feature vector.
[0085] S102d: Determine the classification prediction result corresponding to the training sample according to the content feature vector and the conditional feature vector.
[0086] Optionally, after correlating the content feature vector and the conditional feature vector of the training sample, the result is input into a multi-layer deep neural network (Deep Neural Networks, DNN) to obtain the classification prediction result corresponding to the training sample.
[0087] In the above embodiment, according to the content feature vector representing the sample image and the conditional feature vector representing the classification rule of the sample image, the classification prediction result corresponding to the training sample is determined, so that the determination of each classification prediction result conforms to the search rule of the corresponding sample, improving the accuracy of the prediction result.
[0088] In another possible implementation, the sample resource includes multiple sample images, and the neural network model includes at least an image classification network and a self-attention network. Combining Figure 2 , as Figure 3 shown, S102b includes S102b1 - S102b2.
[0089] In S102b1, multiple sample images are input into the image classification network for feature extraction to obtain the feature vector of each sample image.
[0090] In one implementation, multiple sample images of the training sample are input into the image classification network, such as inception-v3 or resent-50d, for image feature extraction to obtain the feature vector of the sample image. For example, it is pre-set to extract a feature vector of D dimensions, and the number of sample images is M, then a total of M * D-dimensional feature vectors are obtained, that is, M D-dimensional feature vectors.
[0091] In S102b2, the feature vector of each sample image is input into the self-attention network for feature interaction between every two sample images to obtain the content feature vector.
[0092] In one embodiment, the content feature vector can be determined by the reduce_max method. Specifically, the feature vectors of each sample image are input into the self-attention network (i.e., the Transformer network). For example, M D-dimensional feature vectors are input into the Transformer network to learn the feature interaction relationships among the M sample images, obtaining M D-dimensional target feature vectors. Then, the maximum eigenvalue corresponding to each dimension of the M target feature vectors is determined as the content feature vector, resulting in a D-dimensional content feature vector. In another implementation, the average value of each dimension of the M target feature vectors can be determined as the content feature vector, obtaining a D-dimensional content feature vector.
[0093] In the above embodiments, after feature extraction is performed on each sample image, the feature vectors of each obtained sample image are input into the self-attention network to learn the feature interaction relationships among each sample image, which can improve the expression ability of the content feature vector, enabling the feature vector to more accurately represent the features of the sample image.
[0094] In another possible embodiment, S102c includes: determining the target number of rows according to the classification task identifier; determining that the content corresponding to the target number of rows in the pre-set dictionary matrix is the conditional feature vector.
[0095] Optionally, according to the classification task identifier of each resource classification task, the Embedding vector of each resource classification task, that is, the conditional feature vector, is extracted.
[0096] In one embodiment, the dimension of the dictionary matrix is (T, D), where T represents the number of tasks and D represents the dimension of the content feature vector. The classification task identifier of each resource classification task is the sorting serial number of each resource classification task among multiple resource classification tasks. According to the classification task identifier of each resource classification task, a unique D-dimensional vector can be determined in the dictionary matrix, and this D-dimensional feature vector is the Embedding vector of each classification task identifier.
[0097] In the above embodiments, according to the pre-set dictionary matrix and the classification task identifier, the conditional feature vector representing the classification rule of each resource classification task is determined, thereby configuring different classification rules for the sample resources.
[0098] In another possible embodiment, the neural network model includes a deep neural network. S102d includes: performing a multi-head self-attention mechanism on the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; inputting the joint vector into a multi-layer deep neural network to obtain a classification prediction result corresponding to the training sample.
[0099] Optionally, when performing the multi-head self-attention mechanism (i.e., multi-head attention) on the content feature vector and the conditional feature vector, the conditional feature vector is used as the Query, and the content feature vector is used as the Key and Value to implement the conditional probability: P(content feature vector | conditional feature vector), that is, different conditional feature vectors activate different content feature vectors, that is, the conditional feature vector of the training sample activates the content feature vector of the training sample, and then the content feature vector and the conditional feature vector of the training sample are associated.
[0100] In one implementation, a linear transformation is respectively performed on the Query (i.e., the conditional feature vector), the Key (i.e., the content feature vector), and the Value (i.e., the content feature vector), and then input into the scaled dot-product attention (self-attention network), and the above operations are repeated a preset number of times (each time is one head), and then the results of the scaled dot-product attention (self-attention network) for the preset number of times are concatenated, and the value obtained by performing another linear transformation on the concatenated result is used as the result of the multi-head attention (self-attention network), that is, the joint vector of the training sample.
[0101] Further, since the conditional feature vector of the training sample activates the content feature vector of the training sample, after inputting the joint vector of the training sample into the multi-layer deep neural network (Deep Neural Networks, DNN), the multi-layer deep neural network can execute the classification rule represented by the conditional feature vector on the content feature vector of the training sample, and then obtain the classification prediction result corresponding to the training sample.
[0102] In one implementation, as Figure 4As shown below, taking three resource classification tasks as an example of multiple resource classification tasks, preprocessing is first performed on the training sample videos of each resource classification task. For example, global images of N frames are obtained, object detection is performed, and local images of T frames are obtained, resulting in M frames (M = N + T) of sample resources (i.e., M sample images) for each resource classification task. Subsequently, the M-frame sample images and classification task identifiers of each resource classification task are respectively input into different branches of the neural network model. One branch is used to process the M sample images. For example, feature extraction is performed on the M sample images to obtain feature vectors of the M sample images (for example, each feature vector is D-dimensional), namely E1, E2, ……, Em. Then, the feature vectors of the M sample images are input into a self-attention network, such as a Transformer network, to learn the interaction relationships between the M sample images, obtaining M target feature vectors, namely T1, T2, ……, Tm. Finally, the maximum eigenvalue corresponding to each dimension of the M target feature vectors is determined as the content feature vector of the training sample, thereby obtaining a D-dimensional content feature vector. Another branch is used to process the classification task identifiers. For example, the classification task identifier of the first task is 1, the classification task identifier of the second task is 1, and the classification task identifier of the third task is 2, obtaining the conditional feature vectors of the training samples of each resource classification task. Subsequently, a self-attention mechanism is performed on the content feature vector and the conditional feature vector to obtain P(X|C), where X represents the content feature vector and C represents the conditional feature vector, and further obtaining the classification prediction result corresponding to the training sample.
[0103] It should be noted that, as Figure 4 shown below, taking the first task as an example, the sample resources of each resource classification task are obtained through object detection. Among them, the left sample resource is a sample resource with the first content and can be used as a positive sample for the first task, and the right sample resource is a sample resource without the first content and can be used as a negative sample for the first task.
[0104] In the above embodiment, by performing a multi-head self-attention mechanism on the content feature vector and the conditional feature vector, the conditional feature vector of the training sample activates the content feature vector of the training sample, enabling the deep neural network to obtain the classification prediction result by performing the classification rule represented by the conditional feature vector on the content feature vector, resulting in higher prediction accuracy for the task type corresponding to the conditional feature vector. Moreover, the joint vectors of each resource classification task are all input into a deep neural network, that is, multiple resource classification tasks share a deep neural network as the output network, avoiding an overly complex model structure.
[0105] In another possible implementation, the neural network model includes a deep neural network. S102d includes: concatenating the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; and inputting the joint vector into a multi-layer deep neural network to obtain a classification prediction result corresponding to the training sample.
[0106] Optionally, the content feature vector and the conditional feature vector are concatenated to obtain a joint vector of the training sample, so that the content feature vector and the conditional feature vector of the training sample are associated. Then, after inputting the joint vector of the training sample into a multi-layer deep neural network (DNN), the deep neural network can execute the classification rule characterized by the conditional feature vector on the content feature vector of the training sample, and thus obtain the classification prediction result corresponding to the training sample.
[0107] In one implementation, for example, the content feature vector is a D-dimensional feature vector, and the conditional feature vector is also a D-dimensional feature vector. After concatenating the content feature vector and the conditional feature vector of the training sample, a joint feature vector with a dimension of 2*D is obtained.
[0108] In the above embodiment, the content feature vector and the conditional vector are concatenated and then input into the deep neural network, so that the deep neural network can execute the classification rule characterized by the conditional feature vector on the content feature vector, and thus obtain the classification prediction result, making the prediction accuracy for the task type corresponding to the conditional feature vector higher. Moreover, the joint vector of each resource classification task is input into a deep neural network, that is, multiple resource classification tasks share a deep neural network as the output network, avoiding the model structure from being too complex and reducing the complexity of the model structure.
[0109] In another possible implementation, the training method of the resource classification model further includes: performing multi-task joint training on the neural network model according to the training samples of each resource classification task and the training samples of the new resource classification task to obtain a second resource classification model. The training samples of the new resource classification task include the sample resources of the new resource classification task, the classification task identifier of the new resource classification task, and the new label information. The second resource classification model is used to execute multiple resource classification tasks and the new resource classification task.
[0110] The new label information is used to indicate the reference category of the training samples of the new resource classification task.
[0111] In one implementation, after multiple resource classification task models are trained, for example, the training of 3 resource classification tasks has been completed, namely the first resource classification task, the second resource classification task, and the third resource classification task. When it is necessary to add the training of the fourth task, the training samples of the previous three resource classification tasks and the training samples of the newly added resource classification task (for example, the fourth resource classification task) are input into the neural network model for retraining to obtain the second resource classification model. At this time, the second resource classification model can execute the first resource classification task, the second resource classification task, the third resource classification task, and the fourth resource classification task.
[0112] It should be noted that when it is necessary to add multiple resource classification tasks, for example, adding the fifth resource classification task and the sixth resource classification task, the training principle is the same as that when adding one resource classification task, and will not be elaborated here.
[0113] In the above embodiment, when it is necessary to add a resource classification task to the resource classification model that has completed training, the neural network model is retrained jointly with multiple tasks by combining the historical training samples and the training samples of the newly added resource classification task. The obtained resource classification model can possibly execute the newly added resource classification task and the historical resource classification tasks simultaneously. The model has strong scalability, and the newly obtained resource classification model uses the classification prediction results of the newly added resource classification task during model optimization, so the prediction effect of the model is good.
[0114] In another possible implementation, the training method of the resource classification model further includes: obtaining a prediction sample, where the prediction sample includes a prediction sample image and a prediction task identifier; inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction result.
[0115] In one implementation, after the resource classification model is trained, the resource classification model trained with the prediction sample is used for testing to check the test accuracy of the resource classification model. Specifically, the prediction sample image and the identifier of the prediction task to be executed by the prediction sample image (i.e., the prediction task identifier) are input into the resource classification model together to obtain the prediction result obtained after the resource classification model executes the task corresponding to the task identifier for the prediction sample image.
[0116] It should be noted that the prediction task identifier is the classification task identifier in the training samples of any resource classification task. The prediction task identifier is used to represent the task type of the resource classification task executed on the prediction sample image.
[0117] In the above embodiment, when predicting the prediction sample, the prediction sample image and the prediction task identifier are input into the resource classification model simultaneously to obtain the prediction result output by the task model after executing the task corresponding to the task identifier for the prediction sample image. The usage method is simple and the prediction accuracy is high.
[0118] In another possible implementation, inputting the predicted sample image and the prediction task identifier into the first resource classification model to obtain a prediction result, including: inputting the predicted sample image and the prediction task identifier into the first resource classification model to obtain a prediction probability; when the prediction probability is greater than or equal to a preset threshold, determining that the predicted sample is a positive sample of the task type corresponding to the prediction task identifier.
[0119] Optionally, after the first resource classification model outputs a prediction for performing the task type corresponding to the prediction task identifier on the predicted sample image to obtain a prediction probability, when the prediction probability is greater than or equal to a preset threshold (the preset threshold can be 0.75), determining that the predicted sample is a positive sample of the task type corresponding to the prediction task identifier, that is, determining that the prediction result of the predicted sample is "yes". For example, if the task type corresponding to the prediction task identifier is the third task, a prediction result of "yes" indicates that the predicted sample includes the third content.
[0120] In the above embodiment, according to the relationship between the prediction probability output by the resource classification model and the preset threshold, it is determined whether the predicted sample is a positive sample, and thus the binary classification result of the predicted sample can be directly obtained based on the prediction probability, which can simply and intuitively express the prediction result.
[0121] In another possible implementation, S101 includes:
[0122] Step 1: Obtain the training sample video of each resource classification task and the label information corresponding to the training sample video, and determine the classification task identifier of each resource classification task.
[0123] Wherein, the label information corresponding to the training sample video is the label information in the training sample in S101.
[0124] In one implementation, the training sample video of each resource classification task is used to determine the sample resources in the training samples of each resource classification task.
[0125] Optionally, the training sample video includes a positive sample video and a negative sample video. That is, each resource classification task is configured with a positive sample video and a negative sample video. Among them, the positive sample video is used to determine the positive samples of each resource classification task, and the negative samples are used to determine the negative samples of each resource classification task.
[0126] In one implementation, after manually collecting the training sample videos of multiple resource classification tasks, input them into an electronic device, and the electronic device obtains the training sample videos of each resource classification task.
[0127] In another implementation, the electronic device pulls videos from different sources on the network as the training sample videos of each resource classification task.
[0128] Optionally, determining the classification task identifier for each resource classification task can be performed by an electronic device. For example, the electronic device performs one-hot encoding for each resource classification task. For example, it is encoded as (0, 1, …, T–1), where T is a positive integer greater than 1, and determines the encoding of each resource classification task as the classification task identifier.
[0129] Optionally, the classification task identifier for each resource classification task can also be manually input into the electronic device. At this time, the electronic device determines the identifier of each resource classification task input manually as the classification task identifier for that task.
[0130] Perform the following steps two to three for the training sample video of each resource classification task:
[0131] Step two: Obtain the global image of the first preset number of frames of the training sample video.
[0132] Optionally, the electronic device samples the training sample video to obtain the global image of the first preset number of frames.
[0133] In one implementation, each second of the training sample video is determined as one frame, and the first preset number of frames is less than the total number of frames of the training sample video. For example, the training sample video has a total of 20 frames, and the 20-frame training sample video is sampled to obtain N frames of global images. N can be any value greater than 1 and less than 20. For example, N can be equal to 10.
[0134] Step three: Perform object detection on each frame of the global image to obtain the local images of the second preset number of frames. The local images include the object part regions. The global images of the first preset number of frames and the local images of the second preset number of frames are the sample resources for each resource classification task.
[0135] Optionally, the electronic device performs object detection on each frame of the global image to obtain the local images of the third preset number of frames. Further, based on the local images of the third preset number of frames of each frame of the global image, the local images of the second preset number of frames are obtained.
[0136] In one implementation, the electronic device performs object detection on each frame of the global image and determines the detected object part region as the local image of that frame of the global image. Further, when the number of local images including the object part region is less than the third preset number of frames, the cover frame (i.e., the first frame of the video) of the training sample video of this task is used for supplementation. For example, the third preset number of frames is 3, and the local image determined by object detection of the global image is 1. At this time, two cover frames need to be supplemented into the local image set corresponding to this global image.
[0137] It should be noted that the object part regions on each local image are part regions belonging to different objects. For example, there are two objects (object A and object B) on the entire image. At this time, two local images can be obtained according to this global image. Among them, one local image includes the object part region of object A, and the other local image includes the object part region of object B. When the third preset number of frames is 3, a cover frame needs to be used to supplement the local image set corresponding to this global image.
[0138] In the above embodiments, according to the global image of the training sample video and the local image including the object part region, the sample resources are determined, the richness and integrity of the sample resources are improved, and further the expression ability of the content feature vector is improved, making the classification prediction effect more accurate.
[0139] In another possible implementation manner, the training method of the resource classification model further includes: determining the number of rows of the matrix according to the number of tasks corresponding to multiple resource classification tasks; determining the number of columns of the matrix according to the dimension of the preset content feature vector; and obtaining a dictionary matrix according to the number of rows of the matrix, the number of columns of the matrix, and the preset model.
[0140] Optionally, after determining the number of rows and columns of the matrix, input the number of rows and columns of the matrix into the model for generating the dictionary matrix to obtain the dictionary matrix.
[0141] In one implementation manner, each row vector in the dictionary matrix represents the classification rule of a task. Among them, the row vectors of the field matrix correspond one-to-one with the classification task identifiers. That is, the vector in the first row of the dictionary matrix is used to represent the classification rule of the task type corresponding to the classification task identifier 1.
[0142] Optionally, obtaining the dictionary matrix according to the number of rows of the matrix, the number of columns of the matrix, and the preset model includes: obtaining an initial dictionary matrix according to the number of rows of the matrix, the number of columns of the matrix, and the preset model; and randomly initializing the initial dictionary matrix to obtain the dictionary matrix.
[0143] In the above embodiments, the dimension of the dictionary matrix is determined according to the number of tasks and the dimension of the content feature vector, so that the conditional feature vector obtained from the dictionary matrix according to the classification task identifier has the same dimension as the content feature vector, thereby improving the accuracy of the prediction result.
[0144] In another possible implementation manner, the training method of the resource classification model further includes: in each round of iterative training process of the neural network model, updating the dictionary matrix according to the classification prediction result corresponding to the training sample and the label information in the training sample. The updated dictionary matrix is used to determine the conditional feature vector during the next round of iterative training of the neural network model.
[0145] In one embodiment, after the electronic device obtains the classification prediction results corresponding to the training samples of each resource classification task, it determines the loss function values of the training samples of each resource classification task according to the classification prediction results corresponding to the training samples of each resource classification task and the label information in the training samples, and further obtains multiple loss function values of multiple resource classification tasks. Then, the dictionary matrix is updated according to the average value of the multiple loss function values.
[0146] In another embodiment, the electronic device calculates the loss function values by using the classification prediction results corresponding to the training samples of each resource classification task and the label information of the training samples of each resource classification task. For example, the cross-entropy loss function is used to determine the loss function values. Then, the dictionary matrix is updated according to the obtained loss function values.
[0147] In the above embodiments, in each round of iterative training process of the neural network model, the dictionary matrix is updated according to the classification prediction results corresponding to the training samples and the label information in the training samples, so as to update the conditional feature vectors of the classification task identifiers of each resource classification task. Furthermore, the conditional feature vectors can more accurately express the classification rules of the task types corresponding to the classification task identifiers, which helps to improve the accuracy of the classification prediction results corresponding to the training samples, and further improve the training efficiency of the first resource classification model.
[0148] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0149] The embodiments of the present disclosure also provide a training device for a resource classification model.
[0150] Figure 5 is a block diagram of a training device for a resource classification model shown according to an exemplary embodiment. Refer to Figure 5 , the training device 500 for the resource classification model includes an acquisition module 501, a training module 502, an update module 503, and an iteration module 504.
[0151] An acquisition module 501, configured to acquire training samples for each of multiple resource classification tasks; the training samples include sample resources for the corresponding resource classification tasks, classification task identifiers for the corresponding resource classification tasks, and label information for indicating the reference categories of the sample resources in the corresponding training samples. For example, in combination with Figure 1 , the acquisition module 501 can be used to execute S101.
[0152] A training module 502, configured to input the training samples for each of the multiple resource classification tasks into a neural network model to obtain classification prediction results corresponding to the training samples; the classification prediction results are determined according to the content features of the sample resources in the training samples and the classification rules of the task types corresponding to the classification task identifiers. For example, in combination with Figure 1 , the training module 502 can be used to execute S102.
[0153] An update module 503, configured to update the parameters of the neural network model according to the classification prediction results corresponding to the training samples and the label information in the training samples. For example, in combination with Figure 1 , the update module 503 can be used to execute S103.
[0154] An iteration module 504, configured to perform iterative training on the updated neural network model until the neural network model meets the model convergence condition, and determine the converged neural network model as the first resource classification model. For example, in combination with Figure 1 , the iteration module 504 can be used to execute S104.
[0155] In a possible implementation manner, the training module 502 is specifically configured to perform: input the training samples for each resource classification task into the neural network model, and the neural network model performs the following steps: according to the sample resources for each resource classification task, obtain a content feature vector, which is used to characterize the content features of the sample resources for each resource classification task; according to the classification task identifier for each resource classification task, obtain a conditional feature vector, which is used to characterize the classification rules of the task types corresponding to the classification task identifiers of each sample resource task; the dimension of the conditional feature vector is the same as the dimension of the content feature vector; according to the content feature vector and the conditional feature vector, determine the classification prediction results corresponding to the training samples.
[0156] In another possible implementation manner, the sample resources include multiple sample images, and the neural network model includes at least an image classification network and a self-attention network. The training module 502 is specifically configured to perform: input the multiple sample images into the image classification network for feature extraction to obtain a feature vector for each sample image; input the feature vector for each sample image into the self-attention network to perform feature interaction between every two sample images to obtain a content feature vector.
[0157] In another possible implementation, the training module 502 is specifically configured to perform: determining the target number of rows according to the classification task identifier; determining the content corresponding to the target number of rows in the preset dictionary matrix as the conditional feature vector.
[0158] In another possible implementation, the neural network model includes a multi-layer deep neural network, and the training module 502 is specifically configured to perform: performing a multi-head self-attention mechanism on the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; inputting the joint vector into the multi-layer deep neural network to obtain a classification prediction result corresponding to the training sample.
[0159] In another possible implementation, the neural network model includes a multi-layer deep neural network, and the training module 502 is specifically configured to perform: concatenating the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; inputting the joint vector into the multi-layer deep neural network to obtain a classification prediction result corresponding to the training sample.
[0160] In another possible implementation, the training module 502 is further configured to perform: performing multi-task joint training on the neural network model according to the training samples of each resource classification task and the training samples of the new resource classification task to obtain a second resource classification model; wherein, the training samples of the new resource classification task include the sample resources of the new resource classification task, the classification task identifier of the new resource classification task, and the new label information, and the second resource classification model is used to perform multiple resource classification tasks and the new resource classification task.
[0161] In another possible implementation, the device further includes a prediction module, which is configured to perform: obtaining a prediction sample, where the prediction sample includes a prediction sample image and a prediction task identifier; inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction result.
[0162] In another possible implementation, the prediction module is specifically configured to perform: inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction probability; in the case where the prediction probability is greater than or equal to a preset threshold, determining that the prediction sample is a positive sample of the task type corresponding to the prediction task identifier.
[0163] In another possible implementation, the apparatus further includes an acquisition module configured to perform: acquiring a training sample video of each resource classification task and label information corresponding to the training sample video, and determining a classification task identifier for each resource classification task; performing the following steps on the training sample video of each resource classification task: acquiring a global image of a first preset number of frames of the training sample video; performing object detection on each frame of the global image to acquire a local image of a second preset number of frames, the local image including an object part region, and the global image of the first preset number of frames and the local image of the second preset number of frames being sample resources for each resource classification task.
[0164] In another possible implementation, the apparatus further includes a configuration module configured to perform: determining the number of rows of a matrix according to the number of tasks corresponding to multiple resource classification tasks; determining the number of columns of the matrix according to the dimension of a preset content feature vector; and obtaining a dictionary matrix according to the number of rows of the matrix, the number of columns of the matrix, and a preset model.
[0165] In another possible implementation, the configuration module is further configured to perform: in each round of iterative training of the neural network model, updating the dictionary matrix according to the classification prediction result corresponding to the training sample and the label information in the training sample; and using the updated dictionary matrix to determine a conditional feature vector during the next round of iterative training of the neural network model.
[0166] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0167] Figure 6 is a block diagram of an electronic device shown according to an exemplary embodiment. As Figure 6 shown, the electronic device 600 includes but is not limited to: a processor 601 and a memory 602.
[0168] Among them, the above-mentioned memory 602 is used to store executable instructions of the above-mentioned processor 601. It can be understood that the above-mentioned processor 601 is configured to execute instructions to implement the resource classification model training method shown in any one of the above embodiments Figures 1 to 4 in the above embodiments.
[0169] It should be noted that those skilled in the art can understand that Figure 6 the electronic device structure shown in Figure 6 does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than
[0170] The processor 601 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 602, and calling the data stored in the memory 602, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 601 may include one or more processing units; optionally, the processor 601 may integrate an application processor and a modem processor, where the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 601 either.
[0171] The memory 602 can be used to store software programs and various data. The memory 602 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required by at least one functional module (such as a training module, an update module, an iteration module, etc.). In addition, the memory 602 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0172] In an exemplary embodiment, the embodiments of the present disclosure also provide a computer-readable storage medium including instructions, such as the memory 602 including instructions. The above instructions can be executed by the processor 601 of the electronic device 600 to complete the above embodiments Figures 1 to 4 for the training method of the resource classification model shown in any one of the above.
[0173] In actual implementation, the processing functions of the acquisition module 501, the training module 502, the update module 503, and the iteration module 504 can be Figure 6 implemented by the processor 601 shown in calling the program code in the memory 602. The specific execution process can refer to Figures 1 to 4 the description of the training part of the resource classification model shown in any one of the above, which will not be elaborated here.
[0174] Optionally, the computer-readable storage medium can be a non-temporary computer-readable storage medium. For example, the non-temporary computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0175] In an exemplary embodiment, the embodiments of the present disclosure also provide a computer program product including one or more instructions. The one or more instructions can be executed by the processor 601 of the electronic device 600 to complete the above embodiments Figures 1 to 4The training method of the resource classification model shown in any one of the above.
[0176] It should be noted that when one or more instructions in the above instructions in the computer-readable storage medium or the computer program product are executed by the processor 601 of the electronic device 600, each process of the above verification method embodiment is implemented, and the same technical effects as those of the training method of the resource classification model shown in any one of the above embodiments can be achieved. To avoid repetition, it will not be elaborated here. Figures 1 to 4 The same technical effects as those of the training method of the resource classification model shown in any one of the above. To avoid repetition, it will not be elaborated here.
[0177] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0178] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A training method for a resource classification model, characterized in that, Including: Obtaining training samples for each resource classification task among multiple resource classification tasks; The training samples include sample resources for the corresponding resource classification task, a classification task identifier for the corresponding resource classification task, and label information, where the label information is used to indicate the reference category of the training samples for the corresponding resource classification task; among them, the multiple resource classification tasks include multiple tasks with different task types, or the multiple resource classification tasks include multiple tasks with the same task type in different geographical regions; the sample resources include multiple sample images; Inputting the training samples for each resource classification task among the multiple resource classification tasks into a neural network model to obtain classification prediction results corresponding to the training samples; the sample resources and the classification task identifier in the training samples for each resource classification task form a pair group, and different classification task identifiers are used to distinguish the classification tasks of different sample resources, and the classification prediction results are determined according to the content features of the sample resources in the training samples and the classification rules of the task types corresponding to the classification task identifiers; Updating the parameters of the neural network model according to the classification prediction results corresponding to the training samples and the label information in the training samples; Performing iterative training on the updated neural network model until the neural network model meets the model convergence condition, and determining the converged neural network model as the first resource classification model.
2. The method according to claim 1, wherein The step of inputting the training samples for each resource classification task among the multiple resource classification tasks into a neural network model to obtain classification prediction results corresponding to the training samples includes: Inputting the training samples for each resource classification task into the neural network model, and the neural network model performs the following steps: Obtaining a content feature vector according to the sample resources for each resource classification task, where the content feature vector is used to characterize the content features of the sample resources for each resource classification task; Obtaining a conditional feature vector according to the classification task identifier for each resource classification task, where the conditional feature vector is used to characterize the classification rules of the task types corresponding to the classification task identifiers for each resource classification task; the dimension of the conditional feature vector is the same as the dimension of the content feature vector; Determining the classification prediction results corresponding to the training samples according to the content feature vector and the conditional feature vector.
3. The method according to claim 2, wherein The neural network model includes at least an image classification network and a self-attention network. The step of obtaining a content feature vector according to the sample resources for each resource classification task includes: Inputting the multiple sample images into the image classification network for feature extraction to obtain a feature vector for each sample image; Inputting the feature vector for each sample image into the self-attention network to perform feature interaction between every two sample images to obtain the content feature vector.
4. The method according to claim 2, wherein The step of obtaining a conditional feature vector according to the classification task identifier for each resource classification task includes: Determining a target number of rows according to the classification task identifier; Determining that the content corresponding to the target number of rows in a pre-set dictionary matrix is the conditional feature vector.
5. The method according to claim 2, wherein The neural network model includes multiple layers of deep neural networks. Determining the classification prediction result corresponding to the training sample according to the content feature vector and the conditional feature vector includes: Performing a multi-head self-attention mechanism on the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; Inputting the joint vector into the multiple layers of deep neural networks to obtain the classification prediction result corresponding to the training sample.
6. The method according to claim 2, wherein The neural network model includes multiple layers of deep neural networks. Determining the classification prediction result corresponding to the training sample according to the content feature vector and the conditional feature vector includes: Concatenating the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; Inputting the joint vector into the multiple layers of deep neural networks to obtain the classification prediction result corresponding to the training sample.
7. The method according to claim 1, characterized in that The method further includes: Performing multi-task joint training on the neural network model according to the training samples of each resource classification task and the training samples of the new resource classification task to obtain a second resource classification model; wherein, the training samples of the new resource classification task include the sample resources of the new resource classification task, the classification task identifier of the new resource classification task, and new label information, and the second resource classification model is used to perform the multiple resource classification tasks and the new resource classification task.
8. The method according to claim 1, characterized in that, The method further includes: Obtaining a prediction sample, where the prediction sample includes a prediction sample image and a prediction task identifier; Inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction result.
9. The method according to claim 8, wherein The step of inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction result includes: Inputting the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction probability; In the case where the prediction probability is greater than or equal to a preset threshold, determining that the prediction sample is a positive sample of the task type corresponding to the prediction task identifier.
10. The method according to claim 1, characterized in that, The step of obtaining the training samples of each resource classification task among the multiple resource classification tasks includes: Obtaining the training sample video of each resource classification task and the label information corresponding to the training sample video, and determining the classification task identifier of each resource classification task; Performing the following steps on the training sample video of each resource classification task: Obtaining the global image of the first preset number of frames of the training sample video; Performing object detection on each frame of the global image to obtain the local image of the second preset number of frames, where the local image includes an object part area, and the global image of the first preset number of frames and the local image of the second preset number of frames are the sample resources of each resource classification task.
11. The method according to claim 4, wherein The method further includes: Determining the number of rows of the matrix according to the number of tasks corresponding to the multiple resource classification tasks; Determining the number of columns of the matrix according to the dimension of the content feature vector set in advance; Obtaining the dictionary matrix according to the number of rows of the matrix, the number of columns of the matrix, and a preset model.
12. The method according to claim 11, wherein The method further includes: In each round of iterative training of the neural network model, update the dictionary matrix according to the classification prediction result corresponding to the training sample and the label information in the training sample; the updated dictionary matrix is used to determine the conditional feature vector during the next round of iterative training of the neural network model.
13. A training device for a resource classification model, characterized in that, It includes: An acquisition module, configured to acquire training samples for each resource classification task among multiple resource classification tasks; The training samples include sample resources for corresponding resource classification tasks, classification task identifiers for corresponding resource classification tasks, and label information, where the label information is used to indicate the reference category of the training samples for each resource classification task; among them, the multiple resource classification tasks include multiple tasks with different task types, or the multiple resource classification tasks include multiple tasks with the same task type in different geographical regions; the sample resources include multiple sample images; A training module, configured to input the training samples for each resource classification task among the multiple resource classification tasks into a neural network model to obtain the classification prediction result corresponding to the training sample; the sample resources and the classification task identifier in the training samples for each resource classification task form a pair group, and different classification task identifiers are used to distinguish the classification tasks of different sample resources, and the classification prediction result is determined according to the content features of the sample resources in the training sample and the classification rules of the task types corresponding to the classification task identifiers; An update module, configured to update the parameters of the neural network model according to the classification prediction result corresponding to the training sample and the label information in the training sample; An iteration module, configured to perform iterative training on the updated neural network model until the neural network model meets the model convergence condition, and determine the converged neural network model as the first resource classification model.
14. The device according to claim 13, characterized in that, The training module is specifically configured to perform: Input the training samples for each resource classification task into the neural network model, and the neural network model performs the following steps: Obtain a content feature vector according to the sample resources for each resource classification task, where the content feature vector is used to characterize the content features of the sample resources for each resource classification task; Obtain a conditional feature vector according to the classification task identifier for each resource classification task, where the conditional feature vector is used to characterize the classification rules of the task types corresponding to the classification task identifiers for each sample resource task; the dimension of the conditional feature vector is the same as the dimension of the content feature vector; Determine the classification prediction result corresponding to the training sample according to the content feature vector and the conditional feature vector.
15. The device according to claim 14, characterized in that, The neural network model includes at least an image classification network and a self-attention network, and the training module is specifically configured to perform: Input the multiple sample images into the image classification network for feature extraction to obtain a feature vector for each sample image; Input the feature vector for each sample image into the self-attention network to perform feature interaction between every two sample images to obtain the content feature vector.
16. The device according to claim 14, characterized in that, The training module is specifically configured to perform: Determine the target number of rows according to the classification task identifier; Determine that the content corresponding to the target number of rows in the pre-set dictionary matrix is the conditional feature vector.
17. The device according to claim 14, characterized in that, The neural network model includes multiple layers of deep neural networks, and the training module is specifically configured to execute: Perform a multi-head self-attention mechanism on the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; Input the joint vector into the multiple layers of deep neural networks to obtain a classification prediction result corresponding to the training sample.
18. The device according to claim 14, characterized in that, The neural network model includes multiple layers of deep neural networks, and the training module is specifically configured to execute: Concatenate the content feature vector and the conditional feature vector to obtain a joint vector of the training sample; Input the joint vector into the multiple layers of deep neural networks to obtain a classification prediction result corresponding to the training sample.
19. The device according to claim 13, characterized in that The training module is further configured to execute: Perform multi-task joint training on the neural network model according to the training samples of each resource classification task and the training samples of the new resource classification task to obtain a second resource classification model; wherein, the training samples of the new resource classification task include the sample resources of the new resource classification task, the classification task identifier of the new resource classification task, and new label information, and the second resource classification model is used to execute the multiple resource classification tasks and the new resource classification task.
20. The device according to claim 13, characterized in that, The device further includes a prediction module configured to execute: Obtain a prediction sample, where the prediction sample includes a prediction sample image and a prediction task identifier; Input the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction result.
21. The device according to claim 20, wherein The prediction module is specifically configured to execute: Input the prediction sample image and the prediction task identifier into the first resource classification model to obtain a prediction probability; In the case where the prediction probability is greater than or equal to a preset threshold, determine that the prediction sample is a positive sample of the task type corresponding to the prediction task identifier.
22. The device according to claim 13, characterized in that, The acquisition module is specifically configured to execute: Obtain the training sample video of each resource classification task and the label information corresponding to the training sample video, and determine the classification task identifier of each resource classification task; Perform the following steps for the training sample video of each resource classification task: Obtain the global image of the first preset number of frames of the training sample video; Perform object detection on each frame of the global image to obtain local images of the second preset number of frames, where the local images include object part regions, and the global image of the first preset number of frames and the local images of the second preset number of frames are the sample resources of each resource classification task.
23. The device according to claim 16, characterized in that, The device further includes a configuration module configured to execute: Determine the number of matrix rows according to the number of tasks corresponding to the multiple resource classification tasks; Determine the number of matrix columns according to the dimension of the pre-set content feature vector; Obtain the dictionary matrix according to the number of matrix rows, the number of matrix columns, and a preset model.
24. The device according to claim 23, characterized in that, The configuration module is further configured to execute: In each round of iterative training of the neural network model, update the dictionary matrix according to the classification prediction result corresponding to the training sample and the label information in the training sample; the updated dictionary matrix is used to determine the conditional feature vector during the next round of iterative training of the neural network model.
25. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the method according to any one of claims 1 to 12.
26. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 12.
27. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions run on the electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Method and device for identifying relation between objects and electronic system
CN112819011A
Method and device for training classification model and data classification
CN113255824A