Information Classification Method, Multimedia Resource Push Method and Device
The target feature extraction model is generated through pre-training and fine-tuning based on the target graphics and text samples corresponding to the target service, which solves the problem of low training efficiency of the multimedia resource classification model and improves the training efficiency and classification accuracy of the information classification model.
Patent Information
- Application Number
- CN202210511911.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-05-11
AI Technical Summary
In the prior art, the training efficiency of the multimedia resource classification model is low, resulting in low information classification efficiency.
By pre-trained multiple pre-trained task models based on pre-set graphic samples, a target feature extraction model shared by multiple pre-trained task models is obtained, and then a classification model to be trained is determined based on the target feature extraction layer and the pre-set classification layer, and the target graphic samples corresponding to the target service are trained.
The cost of manual annotation is reduced, and the training efficiency and classification accuracy of the information classification model are improved.
Smart Images

Figure CN117115602B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning technology, and in particular, to an information classification method, a multimedia resource pushing method, and an apparatus thereof. Background Art
[0002] Multimedia resources on the Internet are widely sourced and large in number. The quality of these multimedia resources is uneven, and some of them may include multimedia resources that are not suitable for being pushed. By classifying multimedia resources, resource pushing can be performed based on the resource classification results.
[0003] In the prior art, generally, a large number of multimedia resources are manually labeled to obtain corresponding training samples, so that a classification model for multimedia resource classification can be trained based on the training samples. However, due to the large number of multimedia resources, the efficiency of generating training samples through manual labeling is low, which in turn leads to low training efficiency of the classification model and further results in low information classification efficiency. Summary of the Invention
[0004] The technical problem to be solved by this application is to provide an information classification model training method, a multimedia resource pushing method, and an apparatus thereof, which can improve the training efficiency of the classification model and the classification ability of the classification model, and further improve the information classification efficiency and classification accuracy.
[0005] To solve the above technical problem, on the one hand, an embodiment of this application provides an information classification method, including:
[0006] Obtain target information of a target image; the target information is the image information and text information of the target image, or the target information is the image information of the target image;
[0007] Classify the target information based on an information classification model to obtain a first classification result of the target image; the information classification model is obtained by training a to-be-trained classification model based on target graphic and text samples corresponding to a target service; the to-be-trained classification model is obtained based on a target feature extraction model and a preset classification layer; the target feature extraction model is obtained by training multiple pre-trained task models respectively based on preset graphic and text samples, and the target feature extraction model is a shared model of the multiple pre-trained task models.
[0008] On the other hand, an embodiment of this application provides a multimedia resource pushing method, including:
[0009] Determine candidate resource information corresponding to candidate multimedia resources; the candidate resource information is the image information and text information corresponding to the candidate multimedia resources, or the candidate resource information is the image information corresponding to the multimedia resources;
[0010] Classify the candidate resource information based on an information classification model to obtain a second classification result of the candidate multimedia resource; the information classification model is obtained by training a classification model to be trained based on target graphic samples corresponding to a target service; the classification model to be trained is obtained based on a target feature extraction model and a preset classification layer; the target feature extraction model is obtained by separately training a plurality of pre-trained task models based on preset graphic samples, and the target feature extraction model is a shared model of the plurality of pre-trained task models;
[0011] Determine a target multimedia resource from the candidate multimedia resources based on the second classification result;
[0012] Push the target multimedia resource.
[0013] On the other hand, an embodiment of the present application provides an information classification device, including:
[0014] A first acquisition module, configured to acquire target information of a target image; the target information is image information and text information of the target image, or the target information is image information of the target image;
[0015] A first classification module, configured to classify the target information based on an information classification model to obtain a first classification result of the target image; the information classification model is obtained by training a classification model to be trained based on target graphic samples corresponding to a target service; the classification model to be trained is obtained based on a target feature extraction model and a preset classification layer; the target feature extraction model is obtained by separately training a plurality of pre-trained task models based on preset graphic samples, and the target feature extraction model is a shared model of the plurality of pre-trained task models.
[0016] On the other hand, an embodiment of the present application provides a multimedia resource pushing device, including:
[0017] A candidate resource information determination module, configured to determine candidate resource information corresponding to a candidate multimedia resource; the candidate resource information is image information and text information corresponding to the candidate multimedia resource, or the candidate resource information is image information corresponding to the multimedia resource;
[0018] A second classification module, configured to classify the candidate resource information based on an information classification model to obtain a second classification result of the candidate multimedia resource; the information classification model is obtained by training a classification model to be trained based on target text and image samples corresponding to a target service; the classification model to be trained is obtained based on a target feature extraction model and a preset classification layer; the target feature extraction model is obtained by training multiple pre-trained task models respectively based on preset text and image samples, and the target feature extraction model is a shared model of the multiple pre-trained task models;
[0019] A target multimedia resource determination module, configured to determine a target multimedia resource from the candidate multimedia resources based on the second classification result;
[0020] A resource push module, configured to push the target multimedia resource.
[0021] On the other hand, the present application provides an electronic device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the information classification method or the multimedia resource push method as described above.
[0022] On the other hand, the present application provides a computer storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the information classification method or the multimedia resource push method as described above.
[0023] Implementing the embodiments of the present application has the following beneficial effects:
[0024] In this application, a target feature extraction model shared by multiple pre-trained task models is first trained based on a preset graphic and text sample, and then a classification model to be trained is determined based on the target feature extraction layer and the prediction classification layer; the classification model to be trained is trained based on the target graphic and text sample corresponding to the target business to obtain an information classification model corresponding to the target business. That is, in this application, the target feature extraction model is generated through pre-training, and then the target feature extraction model is fine-tuned based on the target graphic and text sample corresponding to the target business to obtain an information classification model that can be used in the target business scenario, thereby reducing the cost of manual annotation and improving the training efficiency of the information classification model; further, the model is trained based on the preset graphic and text sample, so that the trained information classification model has the ability to classify based on graphic and text information, and the graphic and text information can describe the information of the multimedia resource from the image dimension and the text dimension, so that the graphic and text information has a good resource feature representation ability, thereby further improving the classification accuracy; in addition, in this application, the target feature extraction model is determined through the training of multiple pre-trained task models, and the training process of the multiple pre-trained task models is a variety of training constraints on the target feature extraction model, so as to improve the feature extraction ability of the target feature extraction model and further improve the classification ability of the information classification model. Therefore, information classification based on the information classification model in this application can improve the information classification efficiency and classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 is a schematic diagram of the implementation environment provided by the embodiments of the present application;
[0027] Figure 2 is a flowchart of the information classification method provided by the embodiments of the present application;
[0028] Figure 3 is a flowchart of a method for expanding a preset graphic and text sample provided by the embodiments of the present application;
[0029] Figure 4 is a flowchart of the training method for the first task model provided by the embodiments of the present application;
[0030] Figure 5 is a flowchart of the training method for the second task model provided by the embodiments of the present application;
[0031] Figure 6 It is a flowchart of another second task model training method provided by an embodiment of the present application;
[0032] Figure 7 It is a flowchart of a third task model training method provided by an embodiment of the present application;
[0033] Figure 8 It is a flowchart of a method for updating an information classification model provided by an embodiment of the present application;
[0034] Figure 9 It is a flowchart of a resource push method provided by an embodiment of the present application;
[0035] Figure 10 It is a flowchart of a method for determining an information classification result provided by an embodiment of the present application;
[0036] Figure 11 It is a schematic diagram of multiple task models provided by an embodiment of the present application;
[0037] Figure 12 It is a schematic structural diagram of an information classification model provided by an embodiment of the present application;
[0038] Figure 13 It is a schematic diagram of an information classification model training device provided by an embodiment of the present application;
[0039] Figure 14 It is a schematic diagram of a multimedia resource push device provided by an embodiment of the present application;
[0040] Figure 15 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0041] To make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0042] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0043] Please refer to Figure 1 , which shows a schematic diagram of the implementation environment provided by the embodiments of the present application. The implementation environment may include: at least one client 110 and a resource push end 120. The client 110 and the resource push end 120 can communicate data through a network.
[0044] Specifically, the client 110 can send a multimedia resource display request to the resource push end 120. The resource push end 120 can push the corresponding target multimedia resource to the client 110 based on the multimedia resource display request. Among them, before pushing the multimedia resource, the resource push end 120 can classify the candidate multimedia resources, and determine the target push resources that meet the push conditions based on the resource classification results. For the operation of the resource push end 120 to classify the candidate multimedia resources, it can be performed when new candidate multimedia resources appear, that is, perform the resource classification operation on the new candidate multimedia resources.
[0045] The client 110 can communicate with the resource push end 120 based on the browser / server mode (Browser / Server, B / S) or the client / server mode (Client / Server, C / S). The client 110 can include: entity devices such as smart phones, tablet computers, laptop computers, digital assistants, smart wearable devices, vehicle-mounted terminals, servers, etc., and can also include software running on the entity devices, such as application programs, etc. The operating systems running on the client 110 in the embodiments of the present application can include, but are not limited to, the Android system, the IOS system, linux, windows, etc.
[0046] The resource push end 120 and the client 110 can establish a communication connection through wired or wireless means. The resource push end 120 can include an independently operating server, or a distributed server, or a server cluster composed of multiple servers, where the server can be a cloud server.
[0047] To solve the problem of low training efficiency of the classification model in the prior art, which further causes low information classification efficiency, an embodiment of the present application provides an information classification method, and its execution subject can be the above-mentioned resource push end. Please refer to Figure 2 , and the method may include:
[0048] S210. Obtain the target information of the target image; the target information is the image information and text information of the target image, or the target information is the image information of the target image.
[0049] Among them, the target image can be any type of image, and the target image may include image information, text information, etc.; specifically, the target image can be the cover image of a multimedia resource, or the resource content image of a multimedia resource, etc. Taking the cover image of the target image as an example, the image information of the target image is the image information included in the cover image; the text information of the target image is the text information included in the cover image, and the text information can be obtained by performing optical character recognition on the cover image; or the text information of the target image can also be the preset text label information corresponding to the target image, and the embodiment of the present application does not make specific limitations.
[0050] The target information being image information and text information, and the target information being image information respectively correspond to two input modalities of information, so that information classification can be performed based on information of different input modalities, improving the flexibility of information classification.
[0051] S220. Classify the target information based on the information classification model to obtain the first classification result of the target image; the information classification model is obtained by training the classification model to be trained based on the target graphic samples corresponding to the target service; the classification model to be trained is obtained based on the target feature extraction model and the preset classification layer; the target feature extraction model is obtained by respectively training multiple pre-trained task models based on the preset graphic samples, and the target feature extraction model is a shared model of the multiple pre-trained task models.
[0052] The information classification model in this embodiment can be obtained through model pre-training. For the training method of the information classification model, it may include the following steps:
[0053] S2202. Obtain the target feature extraction model; the target feature extraction model is obtained by respectively training multiple pre-trained task models based on the preset graphic samples, and the target feature extraction model is a shared model of the multiple pre-trained task models.
[0054] The target feature extraction model in the embodiments of the present application can be a model that has been trained and is used to extract features from text and image information. Further, the target feature extraction model may include an image feature extraction layer, a text feature extraction layer, and a feature fusion layer. When extracting features through the target feature extraction model, the image information in the text and image information to be extracted is input into the image feature extraction layer to obtain image feature information, the text information in the text and image information to be extracted is input into the text feature extraction layer to obtain text feature information, and the image feature information and the text feature information are input into the feature fusion layer to obtain output features corresponding to the text and image information to be extracted. Specifically, for the image feature extraction layer, the text feature extraction layer, and the feature fusion layer, a Transformer structure can be used, or a convolutional network can also be used for the image feature extraction layer.
[0055] The preset text and image samples may include preset image information and preset text information that matches the preset image information. The preset text information that matches the preset image information can be determined based on the text information included in the preset image information, or can be text information associated with the preset image information, such as the title information corresponding to the image information, or the label information corresponding to the image information, etc. The sources of the preset text and image samples may include:
[0056] 1. Content cover image information and corresponding marking information in the field of information flow distribution; specifically, it may include the cover image of the information flow content and the corresponding title, and a text and image sample pair can be generated based on the cover image and the corresponding title; further, for cover images with similar entity information in the title, they can be considered similar cover images, and thus a text and image sample pair can also be generated based on the title and the corresponding similar cover image. The marking information does not require special manual marking and can be automatically collected from the information flow link.
[0057] 2. Image-text matching pair public datasets in the prior art field, such as CC12M (released by Google Research in February 2021, CVPR published a large dataset Conceptual 12M (CC12M), 12 million image-text data pairs for the training of vision-and-language models), CC3M (Conceptual Captions dataset, a dataset containing image URLs and caption pairs, used for the training and evaluation of machine learning image captioning systems), SBU (the SBU Captions dataset initially regarded image captioning as a retrieval task, containing 1 million picture URLs and title pairs), etc. The images and their original descriptions in the Conceptual Captions dataset come from the network, so they represent a wider range of styles; it can also include data with some classification and object detection result labels, such as Flickr30K and Flickr8K, the COCO dataset. These labeled classification data can be used as part of the pre-training set. Simply put, the labels of these pictures and the pictures themselves form an image-text matching pair.
[0058] 3. Data obtained by crawling Internet data; the entity words corresponding to the content tags are counted through the content distributed by the information flow as retrieval keywords, and data in the fields with a large amount of image data and image descriptions are collected through search applications and vertical websites. The crawled images and the keyword texts of the images are used as an image-text matching pair data.
[0059] Multiple pre-training task models are all trained based on preset image-text samples, and different preprocessing modules and postprocessing modules can be corresponding based on different pre-training tasks. Multiple pre-training task models include the same feature extraction model, and multiple pre-training task models share the model parameters of the feature extraction model. The training process of multiple pre-training task models is a variety of training constraints on the target feature extraction model. The training processes of multiple pre-training task models can be carried out synchronously. During the synchronous training process, the continuously updated model parameters of the feature extraction model are shared among multiple pre-training tasks.
[0060] S2204. Based on the target feature extraction model and the preset classification layer, determine the classification model to be trained.
[0061] In the case of obtaining the target feature extraction model by training multiple pre-training task models, in order to be applicable to the classification scenario, a preset classification layer can be connected after the target feature extraction model, so that the corresponding classification model to be trained can be obtained. The preset classification layer in this embodiment can be a supervised classifier, for example, it can be a Logistic Regression (LR) classifier, a support vector machine, etc.
[0062] S2206. Train the to-be-trained classification model based on the target graphic and text samples corresponding to the target service to obtain an information classification model corresponding to the target service.
[0063] The target feature extraction model can be determined after training multiple pre-trained task models based on a large number of graphic and text samples. Thus, the target feature extraction model is a generalized feature extraction model. To enable the target feature extraction model to be applicable to the target service scenario, the to-be-trained classification model including the target feature extraction model can be trained specifically based on the target graphic and text samples corresponding to the target service, so as to obtain an information classification model capable of classifying the information corresponding to the target service. The number of target service samples is generally less than the number of preset graphic and text samples. Thus, after determining the target feature extraction model based on the preset graphic and text samples, the model can be trained based on a small number of target graphic and text samples corresponding to the target service, and then an information classification model applicable to the target service scenario can be obtained, thereby reducing the requirement for the number of target graphic and text samples corresponding to the target service, and it can be applied to target service scenarios where graphic and text samples are sparse or difficult to obtain.
[0064] In this application, the target feature extraction model is generated through pre-training, and then the target feature extraction model is fine-tuned based on the target graphic and text samples corresponding to the target service to obtain an information classification model that can be used in the target service scenario, thereby reducing the cost of manual annotation and improving the training efficiency of the classification model; further, the model is trained based on the preset graphic and text samples, so that the trained information classification model has the ability to classify based on graphic and text information. The graphic and text information can describe the information of the multimedia resource from the image dimension and the text dimension, so that the graphic and text information has a good resource feature representation ability, thereby further improving the classification accuracy; in addition, in this application, the target feature extraction model is determined through the training of multiple pre-trained task models. The training processes of the multiple pre-trained task models are various training constraints on the target feature extraction model, so as to improve the feature extraction ability of the target feature extraction model and further improve the classification ability of the information classification model. Thus, information classification based on the information classification model in this application can improve the information classification efficiency and classification accuracy.
[0065] The preset graphic and text samples include multiple preset graphic and text matching items, and the image information and text information in each preset graphic and text matching item match; to further expand the preset graphic and text samples, for each preset graphic and text matching item in the preset graphic and text samples, the corresponding similar graphic and text matching items can be determined, and the similar graphic and text matching items can be obtained based on the image information and text information in the preset graphic and text samples.
[0066] For details, please refer to Figure 3, which shows a method for expanding a preset graphic and text sample. The method may include:
[0067] S310. Extract information from the text information in the multiple preset graphic and text matching items respectively to obtain named entity information corresponding to the multiple preset graphic and text matching items respectively.
[0068] S320. Based on the named entity information corresponding to each preset graphic and text matching item and the named entity information corresponding to the remaining preset graphic and text matching items, determine the similar graphic and text matching items for each preset graphic and text matching item; the remaining preset graphic and text matching items are the preset graphic and text matching items in the multiple preset graphic and text matching items except each preset graphic and text matching item.
[0069] S330. Generate a first new matching item based on the image information in each preset graphic and text matching item and the text information in the similar graphic and text matching items.
[0070] S340. Generate a second new matching item based on the image information in the similar graphic and text matching items and the text information in each preset graphic and text matching item.
[0071] The named entities in this embodiment may include noun information such as person names, locations, organizations, times, objects, events, etc., so the corresponding named entity information may include one or more of the above named entities. For the extraction of named entity information, it can be obtained through a preset named entity recognition model. When the named entity information corresponding to each preset graphic and text matching item is recognized, the similarity between each preset graphic and text matching item and other preset graphic and text matching items can be calculated based on the named entity information corresponding to each preset graphic and text matching item, so as to obtain the similarity between each preset graphic and text matching item and other preset graphic and text matching items, and then determine the similar graphic and text matching items corresponding to each preset graphic and text matching item. When determining the similar graphic and text matching items corresponding to each preset graphic and text matching item, the preset graphic and text matching items with a similarity greater than the preset similarity can be determined as the similar graphic and text matching items, or the preset number of preset graphic and text matching items with the largest similarity can be determined as the similar graphic and text matching items.
[0072] The image information corresponding to text information with similar named entity information can also be considered similar. Thus, when the similar graphic-text matching item corresponding to the current preset graphic-text matching item is determined, a first new matching item can be formed based on the image information of the current preset graphic-text matching item and the text information in the similar graphic-text matching item. When there are multiple similar graphic-text matching items, a first new matching item can also be formed based on the image information of the current preset graphic-text matching item and the text information in randomly selected multiple similar graphic-text matching items. Further, when the similar graphic-text matching item corresponding to the current preset graphic-text matching item is determined, a second new matching item can also be generated based on the image information in the similar graphic-text matching item and the text information in the current preset graphic-text matching item. When there are multiple similar graphic-text matching items, a second new matching item can also be formed based on the text information of the current preset graphic-text matching item and the image information in randomly selected multiple similar graphic-text matching items.
[0073] In one example, for multiple similar graphic-text matching items corresponding to each preset graphic-text matching item, the image information and text information in the multiple similar graphic-text matching items can be mutually replaced, so as to generate a third new matching item.
[0074] By determining the similar graphic-text matching items for each preset graphic-text matching item and performing mutual replacement of the image information and text information between each preset graphic-text matching item and the corresponding similar graphic-text matching item, new matching items are formed, realizing the expansion of the preset graphic-text samples and increasing the sample quantity.
[0075] The multiple pre-training task models in the embodiments of this application may include a first task model. The first task model may be a model for implementing contrast task learning. By training the first task model, the first task model is enabled to have the performance of recognizing similarities or differences. Specifically, the first task model may include an initial feature extraction model and a contrast output layer. For details, please refer to Figure 4 which shows the training method of the first task model. The method may include:
[0076] S410. Respectively perform feature extraction on the preset graphic-text sample and the contrast graphic-text sample corresponding to the preset graphic-text sample based on the initial feature extraction model to obtain corresponding first output features and second output features. The contrast graphic-text sample is obtained by performing information transformation on the image information and / or text information in the preset graphic-text sample.
[0077] S420. Perform contrast processing on the first output feature and the second output feature based on the contrast output layer to obtain contrast output information.
[0078] S430. Adjust the model parameters of the first task model based on the comparison output information and the preset matching information between the preset graphic and text sample and the comparison graphic and text sample, to obtain the trained first task model.
[0079] S440. Determine the target feature extraction model based on the trained first task model.
[0080] In this embodiment, the preset image information can be subjected to image transformation to obtain transformed image information, and the preset image information is similar or matched with the transformed image information; the preset text information can be subjected to image transformation to obtain transformed text information, and the preset text information is similar or matched with the transformed text information. The comparison graphic and text sample corresponding to the preset graphic and text sample can be a graphic and text sample obtained by transforming the image information in the preset graphic and text sample, or a graphic and text sample obtained by transforming the text information in the preset graphic and text sample, or a graphic and text sample obtained by transforming both the image information and the text information in the preset graphic and text sample. Specifically, the specific methods for performing image transformation on the preset image information may include transformation operations such as rotation, cropping, Gaussian noise, masking, color transformation, and filters, and the specific methods for performing text transformation on the preset text information may include transformation operations such as back translation, character insertion, and deletion.
[0081] Further, in the multimedia resource scenario, the preset image information corresponding to the multimedia resource can be the image information corresponding to any image frame in the multimedia resource, so that the other image frames in the multimedia resource can all be used as the transformed image information of the preset image information. For example, for a video content, it can be considered that adjacent extracted frames are similar, and the extracted frames of different videos and non-adjacent video frames after deduplication are not similar: specifically, negative sample pairs are randomly constructed from the extracted frame times of different videos, and the existing video deduplication relationship chain is used during the construction to avoid duplicate videos. For the information flow content library, the enabled videos in the information content library are used, and adjacent / close frames are extracted from each video and regarded as similar images as positive samples.
[0082] The initial feature extraction model corresponds to the target feature extraction model. After training the initial feature extraction model, the corresponding target feature extraction model can be obtained.
[0083] When performing feature extraction on the preset graphic and text samples and the comparison graphic and text samples based on the initial feature extraction model, it can be implemented based on the initial feature extraction model in the same state. The same state can refer to the state where the model parameters of the initial feature extraction model are the same when performing feature extraction twice, so as to ensure that feature extraction is performed separately in the same state. For example, when performing feature extraction on the preset graphic and text samples, the model parameters of the initial feature extraction model at this time can be recorded, which is convenient for using the same model parameters when performing feature extraction on the comparison graphic and text samples. Further, a siamese network model corresponding to the initial feature extraction model can also be used to perform corresponding feature extraction. The initial feature extraction model and the siamese network model share parameters, thus ensuring the consistency of the input state.
[0084] The comparison output layer can perform comparative analysis on the input first output feature and the second output feature to determine the comparison output information of the two. For the preset graphic and text samples and the transformed graphic and text information, the preset matching information can be determined based on their matching relationship. The preset matching information can be matching or not matching. Based on the comparison output information and the preset matching information, the first loss information can be determined, and then the model parameters of the first task model can be adjusted based on the first loss information to obtain the corresponding trained first task model. The determination of the first loss information can be implemented based on the Contrastive Loss function.
[0085] The first task model includes an initial feature extraction model, and the corresponding trained first task model includes a trained feature extraction model. For the determination of the target feature extraction model, it can be jointly determined based on the trained feature extraction model obtained by training the first task model and the trained feature models obtained by training one or more other task models.
[0086] By the first task model executing the contrastive learning task, the first task model can be enabled to identify the similarities or differences between information, thereby improving the information expression ability of the first task model, and thus also improving the information expression ability of the target feature extraction model determined based on the first task model.
[0087] The pre-training task model in this embodiment may further include a second task model. The second task model can be a model for performing an image-text matching task. By training the second task model, the second task model can learn the correlation between image information and text information. Correspondingly, please refer to Figure 5 , which shows the training method of the second task model. The method may include:
[0088] S510. Obtain the replacement graphic and text matching items corresponding to the multiple preset graphic and text matching items respectively; the replacement graphic and text matching items are obtained by performing information replacement on the text information in the multiple preset graphic and text matching items.
[0089] S520. Divide the matching items based on the multiple preset graphic-text matching items and the replacement graphic-text matching items corresponding to each of the multiple preset graphic-text matching items, to obtain a first replacement item and a second replacement item; the first replacement item includes replacement items where the image information matches the text information and replacement items where the image information does not match the text information; the second replacement item is a replacement item where the image information does not match the text information.
[0090] S530. Train the second task model based on the first replacement item to obtain a trained second task model.
[0091] S540. Perform matching prediction on the second replacement item based on the trained second task model to obtain matching prediction information.
[0092] S550. When the matching prediction information indicates that the image information in the second replacement item matches the text information, determine the second replacement item as the target negative sample.
[0093] S560. Train the trained second task model based on the target negative sample to obtain an updated second task model.
[0094] S570. Determine the target feature extraction model based on the updated second task model.
[0095] When replacing the information in the preset graphic-text matching item, the image information can be replaced, or the text information can be replaced. In this embodiment, taking the replacement of text information as an example, the text information can be replaced with a preset probability, and the preset probability can be 0.5, 1, etc. When replacing the text information in the preset graphic-text matching item, the text information of other preset graphic-text matching items in the preset preset graphic sample can be used to replace the text information in the current preset graphic-text matching item, or the text information in the current preset graphic-text matching item can be replaced with empty information. Replacing the text information in the current preset graphic-text matching item with the text information of other preset graphic-text matching items enables the second task model to learn the correlation between the image information and the text information; by replacing the text information in the current preset graphic-text matching item with empty information, the situation of text information loss in the actual application scenario can be simulated, thereby improving the adaptability and robustness of the second task model.
[0096] In the case of replacing the text information in the current preset graphic-text matching item with the text information of other preset graphic-text matching items in the preset preset graphic sample, the replacement graphic-text matching item may also include replacement items where the image information matches the text information and replacement items where the image information does not match the text information.
[0097] The image information and text information in multiple preset graphic-text matching items match each other, so that multiple preset graphic-text matching items can be directly determined as the first replacement items; for the replacement graphic-text matching items corresponding to each of the multiple preset graphic-text matching items, the replacement items in which the image information and text information match in the replacement graphic-text matching items can be determined as the first replacement items, and some replacement graphic-text matching items are randomly selected from the replacement items in which the image information and text information do not match as the first replacement items, and the remaining replacement graphic-text matching items are determined as the second replacement items.
[0098] The first replacement items include replacement items in which the image information and text information match. The second task model can be trained based on the first replacement items to obtain a trained second model; then, the second replacement items are subjected to matching prediction based on the trained second task model to obtain matching prediction information; the image information and text information in the second replacement items do not match, so the second replacement items can be used as negative samples. The preset matching information of each replacement graphic-text information in the second replacement items is not matching. Therefore, when the matching prediction information indicates that the image information and text information in the second replacement items match, it can be determined at this time that the prediction is incorrect, and the second replacement items in which the image information and text information match are determined as target negative samples. The target negative samples are difficult negative samples, that is, negative samples that are prone to prediction errors. Therefore, the negative samples that are prone to prediction errors are screened out, and the second task model is trained based on the target negative samples, so that the second task model has the ability to learn the target negative samples, and the model performance of the second task model can be improved; further, a target feature extraction model is determined based on the updated second task model, and the model performance of the target feature extraction model is improved.
[0099] For the determination of the target feature extraction model, the target feature extraction model can be jointly determined based on the trained feature extraction model obtained by training the second task model and the trained feature models obtained by training one or more other task models.
[0100] Further, the second task model includes an initial feature extraction model and a matching output layer; correspondingly, please refer to Figure 6 , which shows another second task model training method, including:
[0101] S610. Feature extraction is performed on the first replacement items based on the initial feature extraction model to obtain a third output feature.
[0102] S620. Matching processing is performed on the third output feature based on the matching output layer to obtain matching output information.
[0103] S630. Based on the matching output information and the preset matching information of the image information and text information in the first replacement items, the model parameters of the second task model are adjusted to obtain the trained second task model.
[0104] In this embodiment, the third output feature may include the feature corresponding to the image information of the first replacement item and the feature corresponding to the text information. Accordingly, the matching output layer may perform matching processing on the input third output feature to obtain the matching output information of the image information and the text information in the first replacement item; the matching output information may be used to characterize the matching degree between the image information and the text information in the first replacement item.
[0105] The matching output layer may specifically be a linear ITM head, which may map the input third output feature into binary logits, and determine whether the image information and the text information match based on the binary logits.
[0106] The image information and the text information in the first replacement item have corresponding preset matching information. Based on the matching output information and the preset matching information, the second loss information can be determined, and the model parameters of the second task model can be adjusted based on the second loss information to obtain the corresponding trained second task model.
[0107] By executing the image-text matching learning task through the second task model, the second task model can learn the correlation between the image information and the text information, thereby improving the model performance of the second task model, and thus also improving the model performance of the target extraction model determined based on the second task model.
[0108] The pre-training task model in this embodiment may further include a third task model. The third task model may be a model for implementing a supervised classification training task. By training the third task model, the third task model can learn the classification information contained in the image information and the text information; the third task model includes an initial feature extraction model and a classification output layer; the preset image-text sample includes a plurality of preset image-text matching items, and the image information and the text information in each preset image-text matching item match. Accordingly, please refer to Figure 7 , which shows the training method of the third task model. The method may include:
[0109] S710. Extract features from the null information items corresponding to the plurality of preset image-text matching items based on the initial feature extraction model to obtain a fourth output feature; the null information items are obtained by performing a null processing on the text information in the plurality of preset image-text matching items.
[0110] S720. Perform classification processing on the fourth output feature based on the classification output layer to obtain classification output information.
[0111] S730. Adjust the model parameters of the third task model based on the classification output information and the target classification labels of the plurality of preset image-text matching items to obtain a trained third task model.
[0112] Determine the target feature extraction model based on the trained third task model.
[0113] The classification output layer in the third task model can output corresponding classification information based on the input features. The classification output layer can be a Logistic Regression (LR) classifier, a support vector machine, etc.; the preset graphic and text samples can also include target classification labels corresponding to multiple preset graphic and text matching items respectively. Training the third task model based on multiple preset graphic and text matching items and the corresponding target classification labels can obtain the corresponding trained third task model.
[0114] Further, random nulling processing can be performed on the text information of multiple preset graphic and text matching items. Specifically, some preset graphic and text matching items can be randomly selected from multiple preset graphic and text matching items, and then the selected part of the preset graphic and text matching items are nulled with a preset probability to obtain corresponding nulled information items. The preset probability can be 0.5, 1, etc.; then the third task model can be trained based on the unselected preset graphic and text matching items and the nulled information items.
[0115] Among them, when training the third task model based on the nulled information items, the nulled information items are feature-extracted based on the initial feature extraction model to obtain corresponding fourth output features; the fourth output features are classified based on the classification output layer to obtain classification output information, the third loss information is determined based on the classification output information and the target classification labels, and the model parameters of the third task model are adjusted based on the third loss information to obtain the corresponding trained third task model.
[0116] For the preset graphic and text matching items that are not nulled, the information input into the initial feature extraction model includes image information and text information; for the nulled information items, the information input into the initial feature extraction model includes image information and null information, that is, the image information and text information of the preset graphic and text matching items that are not nulled, and the input forms of the image information and null information of the nulled information items are the same; the same input form can specifically be that the input sequence lengths are the same, and the input information formats are the same, etc.
[0117] For the determination of the target feature extraction model, the target feature extraction model can be jointly determined based on the trained feature extraction model obtained by training the third task model and the trained feature models obtained by training one or more other task models.
[0118] Perform a supervised classification training task through the third task model, enabling the third task model to learn the classification information contained in the image information and text information, thereby improving the model performance of the third task model, and thus also improving the model performance of the target extraction model determined based on the third task model; further, by randomly nullifying the text information in the preset text-image matching item, it is possible to simulate information classification based on the image information in the case of missing text information, thereby improving the adaptability and robustness of the trained third task model, and further improving the adaptability and robustness of the target extraction model determined based on the third task model.
[0119] Further, the pre-training task model in the embodiment of the present application may further include a fourth task model. The fourth task model may be a masked language model. The fourth task model includes an initial feature extraction model and a masked information prediction layer. By masking the text information in the preset text-image matching item, and then based on the initial feature extraction model, extracting features from the image information and the masked text information in the preset text-image matching item to obtain a fifth output feature; predicting the masked prediction information based on the masked information prediction layer for the fifth output feature, and training the fourth task model based on the masked prediction information and the masked text information to obtain a trained fourth task model; determining a target feature extraction model based on the trained fourth task model.
[0120] In this embodiment, the target text-image samples corresponding to the target service may include multiple preset text-image matching items and corresponding target classification labels, so that the classification model to be trained can be trained based on the target text-image samples to obtain a corresponding information classification model; in the target service scenario, according to the service requirements, the number of classifications is increased, so that on the basis of the target classification labels, updated classification labels corresponding to the newly added classifications are added; for details, please refer to Figure 8 , which shows a method for updating an information classification model. The method may include:
[0121] S810. Obtain updated text-image samples; the updated classification labels in the updated text-image samples are different from the multiple target classification labels.
[0122] S820. Perform model training on the information classification model based on the updated text-image samples to obtain an updated information classification model.
[0123] In one example, M target classification labels can be determined in the target business scenario, that is, for the graphic and text information involved in the target business scenario, there can be M types of classifications; as the target business progresses, the classification of the graphic and text information in the target business scenario is updated; in this embodiment, updating the target classification labels can include adding new classification labels on the basis of the target classification labels, thereby increasing the types of classifications; updating the target classification labels can also be to further refine one or more of the target classification labels, that is, each target classification label can be refined into two or more new labels, and these two or more new labels can replace the corresponding target classification labels; when the classification labels are updated, it is also necessary to update the graphic and text information corresponding to the updated classification labels. Thus, a fine-grained division of the classification labels is achieved.
[0124] Thus, model training is performed on the information classification model based on the updated graphic and text sample information, that is, model training after updating the graphic and text samples is achieved on the basis of the already trained information classification model, and there is no need to perform model training again. The re-training of the information classification model can be based on the updated information, so that the information classification model can learn the updated information, thereby improving the flexibility of updating the information classification model and making the information classification model adapt to the actual application scenario of the target business.
[0125] After training the information classification model, information classification can be performed based on the information classification model; specifically, in the resource push scenario, please refer to Figure 9 , which shows a resource push method, and the method can include:
[0126] S910. Determine the candidate resource information corresponding to the candidate multimedia resource; the candidate resource information is the image information and text information corresponding to the candidate multimedia resource, or the candidate resource information is the image information corresponding to the multimedia resource.
[0127] S920. Classify the candidate resource information based on the information classification model to obtain the second classification result of the candidate multimedia resource; the information classification model is obtained by training the classification model to be trained based on the target graphic and text samples corresponding to the target business; the classification model to be trained is obtained based on the target feature extraction model and the preset classification layer; the target feature extraction model is obtained by performing model training on multiple pre-trained task models respectively based on the preset graphic and text samples, and the target feature extraction model is a shared model of the multiple pre-trained task models.
[0128] S930. Determine the target multimedia resource from the candidate multimedia resources based on the second classification result.
[0129] S940. Push the target multimedia resource.
[0130] The multimedia resources in this embodiment can be image-text resources, image-video resources, text-video resources, etc. Thus, the resource information of the multimedia resources can be the image information and text information corresponding to the multimedia resources, or the image information corresponding to the multimedia resources. Specifically, the image information corresponding to the multimedia resources can be the cover image information when displaying the multimedia resources, or the image content information included in the multimedia resources, and the text information corresponding to the multimedia resources can be the title information of the multimedia resources, or the text information identified from the multimedia resources.
[0131] Classify based on the candidate resource information to obtain an information classification result. Here, the information classification result can be the classification result of the candidate resource information, or the classification result of the candidate multimedia resource corresponding to the candidate resource information.
[0132] Determine a target multimedia resource that meets the push conditions from the candidate multimedia resources according to the information classification result, and perform resource push based on the target multimedia resource.
[0133] Since the information classification model is obtained based on the above model training method of this embodiment, the training processes of multiple pre-training task models are various training constraints on the target feature extraction model, which can improve the feature extraction ability of the target feature extraction model and further improve the classification ability of the information classification model; furthermore, it can improve the accuracy and rationality of the target media resource push.
[0134] In this embodiment, information classification can be performed based on multimodal information. The multimodal can include an image-text input modality and a text input modality; correspondingly, please refer to Figure 10 , which shows a method for determining an information classification result. The method includes:
[0135] S1010. When the candidate resource information is the image information corresponding to the candidate multimedia resource, perform a nulling process on the text information corresponding to the candidate multimedia resource to obtain nulling information.
[0136] S1020. Input the image information corresponding to the candidate multimedia resource and the nulling information into the information classification model to obtain the second classification result; the information classification model is obtained by replacing or nulling the text information in the preset graphic and text samples.
[0137] When classifying based on an information classification model, there may be a situation where text information is missing. During the above model training process in this embodiment, by performing a nulling operation or a replacement operation on the text information, the information classification model can adapt to the situation of missing text information and still be able to classify image information in the case of missing text information to obtain corresponding information classification results.
[0138] Thus, the information classification model can adapt to different input modalities, improving the adaptability and flexibility of use of the information classification model.
[0139] The following uses a specific example to illustrate the specific implementation process of this application. Please refer to Figure 11 , which respectively shows the schematic diagrams of the first task model, the second task model, the third task model, and the fourth task model. As can be seen from Figure 11 , the first task model, the second task model, the third task model, and the fourth task model share a feature extraction model.
[0140] Regarding the structure of the information classification model, please specifically refer to Figure 12 . In an inappropriate information recognition scenario, the inappropriate information can be recognized based on the information classification model. The inappropriate information in this embodiment can be information that is sensitive to the viewing object and likely to cause discomfort to the viewing object. Correspondingly, the target graphic and text samples can be collected based on the inappropriate information actively feedback by the viewing object. By fine-tuning the classification model to be trained based on the inappropriate samples, an information classification model for classifying inappropriate information can be obtained. The inappropriate information can specifically include articles, picture sets, videos, etc. Specifically, the information classification model can predict multiple inappropriate types and normal types, and the inappropriate types can also be updated as the business is adjusted. For example, the label types that the information classification model can predict include inappropriate type a, inappropriate type b, inappropriate type c, and normal type. Now, the label types can be updated, which can be the addition of label types. The updated label types include inappropriate type a, inappropriate type b, inappropriate type c, inappropriate type d, inappropriate type e, and normal type; it can also be the refinement of the existing label types. The updated label types include inappropriate type a1, inappropriate type a2, inappropriate type b, inappropriate type c, and normal type.
[0141] Due to the characteristics of discomfort information such as a wide variety of types, complex scenarios, and low proportion, on the one hand, in this embodiment, the discomfort information can be characterized based on the image information and text information of the discomfort information, thereby improving the accuracy and comprehensiveness of the feature expression of the discomfort information, facilitating the information classification model to extract corresponding features, and performing information classification based on the extracted features; on the other hand, by introducing a large-scale pre-trained model, the pre-trained model is fine-tuned based on a small number of discomfort information samples on the basis of the pre-trained model to obtain an information classification model for classifying the discomfort information, thereby improving the accuracy and recall rate of the information classification model.
[0142] In a specific scenario, the classification of discomfort information is determined based on the feedback information of most viewing objects. However, different viewing objects have different acceptance levels of discomfort information. That is, for a small number of viewing objects, the relevant discomfort information can be accepted. When the target information is predicted to be of the discomfort type, the acceptance level of the target information by the information to be pushed can be further determined based on the object portrait information of the object to be pushed; when the acceptance level is greater than the preset value, the target information can be pushed to the object to be pushed. The acceptance level can be determined by matching information based on the object portrait information and the target information.
[0143] The multi-classification results and the probabilities of classification are used for subsequent product strategies, such as filtering and directly not enabling or reducing the weight for distribution according to the probability size. Here, due to the use of large-scale multi-modal pre-training learning and fine-tuning through a small part of the sample data of the actual business tasks, the demand for discomfort picture recognition data samples is effectively reduced, the cost of manually labeled samples is reduced, and the R & D speed and efficiency of the business algorithm process are improved.
[0144] The following focuses on describing the main functions of each service module of the discomfort picture content recognition method and system based on the multi-modal pre-trained model as follows:
[0145] I. PGC and UGC content production and consumption ends
[0146] (1) Content producers of PGC or UGC, MCN or PUGC provide local or captured video content, self-media articles or picture sets written through the mobile or backend interface API system. The author can choose to actively upload the cover picture of the corresponding content, which are the main content sources of information flow distribution content;
[0147] (2) Through communication with the upstream and downstream content interface services, first obtain the upload server interface address, and then upload the local file. During the shooting process, the local video content can be paired with music, filter templates, video beautification functions, etc.;
[0148] (3) As a consumer, it communicates with the content distribution export server to obtain the index information of the corresponding content, that is, the access quality of the content. For video, it communicates with the video storage server, downloads the corresponding streaming media file and plays it through the local player. For pictures and texts, it usually communicates directly with the CDN service deployed at the edge;
[0149] (4) Report the user's browsing behavior data, reading speed, completion rate, reading time, freeze, loading time, playback clicks, etc. during the upload and download process to the server;
[0150] (5) Consumers usually browse consumption data through the Feeds stream, provide a direct reporting and feedback portal for inappropriate image content, and directly connect to the manual review system for confirmation and review. The review results are stored in the inappropriate image content sample library as a small sample data source for subsequent training of business models;
[0151] 2. Uplink and Downlink Content Interface Server
[0152] (1) Communicate directly with the content production end. The content submitted by the front end, usually the title, publisher, summary, cover image, release time, or the video, directly enters the server end through the server and stores the file in the video content storage service;
[0153] (2) Writing metadata of the video content, such as video file size, cover image link, bit rate, file format, title, release time, author, etc., into the content database;
[0154] (3) Submit the uploaded files and content metadata to the dispatch center service for subsequent content processing and circulation;
[0155] 3. Content Database
[0156] (1) The core database of content. The metadata of all content released by producers is stored in this business database, with the focus on the metadata of the content itself, such as file size, cover image link, bit rate, file format, title, release time, author, video file size, video format, whether it is original or first released, and the classification of content during the manual review process;
[0157] (2) During the manual review process, the information in the content database will be read, and the results and status of the manual review will also be sent back to the content database;
[0158] (3) The content processing by the dispatching center mainly includes machine processing and manual review processing. Here, the core of machine processing involves various quality judgments such as low-quality filtering, content tagging such as classification, tag information, and content deduplication. Their results will be written into the content database, and exactly the same duplicate content will not be processed manually for a second time;
[0159] (4) During the process of constructing the samples of pre-trained graphic and text data, it is done by reading from the content data or through the meta-information of the content such as the title, classification, and tags;
[0160] IV. Dispatching Center Services
[0161] (1) Responsible for the entire dispatching process of video and graphic content circulation. For the content stored in the database through the upstream and downstream content interface servers, the meta-information of the content is then obtained from the content meta-information database;
[0162] (2) As the actual dispatching controller for the graphic and video link, according to the type of content, for the picture content in the link, dispatch the multi-modal inappropriate picture content recognition service system to process the corresponding content, directly filter and mark the corresponding content on the content, and use it with reduced weight by the recommendation engine or for personalized targeted distribution;
[0163] (3) Dispatch the manual review system and the machine processing system, and control the order and priority of dispatching;
[0164] (4) Through the manual review system, the content is enabled, and then through the content export distribution service (usually the recommendation engine or search engine or operation), it is directly provided to the content consumers at the terminal on the display page, that is, the content index (such as the URL of the content access entrance with low quality) information obtained by the consumer side;
[0165] V. Manual Review Service and Report of Inappropriate Picture Content for Complaint
[0166] (1) Usually a WEB system, on the link, it undertakes the results of machine filtering, conducts manual confirmation and review of the results, writes down the reviewed results in the content information meta-database, and at the same time, the actual effects of the machine and the filtering model can be evaluated online through the results of this manual review;
[0167] (2) Report the source of receiving tasks in the manual review process, review results, start and end times of review, and other detailed review processes to the statistical server;
[0168] (3) Interface with the complaint and content reporting review system on the user consumption side, prioritize the handling of inappropriate image content in complaints and reports, and after confirmation, directly take effect on similar content in the enabled content in the inappropriate image content library, mainly through content vectorization matching. At the same time, the review results provide a small sample data basis for building an inappropriate image business model based on a multi-modal pre-trained model in the inappropriate image content library.
[0169] VI. Content Storage Service
[0170] (1) Usually a group of storage servers with a wide distribution range and close to the C-side users are accessed nearby. There are usually CDN acceleration servers for distributed caching acceleration on the periphery. The video and image content uploaded by content producers is saved through the upstream and downstream content interface servers.
[0171] (2) After obtaining the content index information, end consumers can also directly access the video content storage server to download the corresponding content.
[0172] (3) In addition to being a data source for external services, it also serves as a data source for internal services, providing the download file system with the original video data for relevant processing. The access paths for internal and external data sources are usually deployed separately to avoid mutual influence.
[0173] VII. Inappropriate Image Sample Library
[0174] (1) Obtain the content marked by manual review from the content meta-information and repository as the prototype small sample data for establishing inappropriate image content.
[0175] (2) Regularly (usually on a daily basis), retrieve inappropriate image content.
[0176] VIII. Multi-modal Image Large-scale Pre-trained Model
[0177] (1) According to the detailed process described above, increasing the representativeness of data generalization through multi-source data can provide much help in reducing the sample requirements for subsequent actual tasks. The main large-scale data sources include data pushed in the information flow content link, publicly available datasets in the technical field, and data crawled from the public Internet through a crawler system.
[0178] (2) Large-scale pre-training of images uses unsupervised (self-supervised contrastive learning) learning and weakly supervised (text-image pair matching) data for training, mainly by constructing different types of pre-trained models, as shown above to build the pre-trained model.
[0179] IX. Inappropriate Image Content Recognition Model and Service
[0180] (1) Based on the above-mentioned multi-modal image large-scale pre-training model, use the small samples in the inappropriate image small sample library to construct an inappropriate image recognition model through model fine-tuning, and then service the model;
[0181] (2) Communicate with the content scheduling center service to construct a service that can be called on the main link of the information flow content transfer to filter inappropriate image content, or mark it to implement subsequent recommended downgraded distribution or targeted recommendation;
[0182] X. Download File System
[0183] (1) Download and obtain the original video content from the content storage server, and control the download speed and progress. Usually, it is a group of parallel servers composed of relevant task scheduling and distribution clusters;
[0184] (2) For the downloaded file, call the frame extraction service to obtain the necessary key frames of the video file from the video source file, which can be used as the input for subsequent construction of video fingerprints or as a candidate source for cover images. By extracting the OCR text in the image, the image can be paired with the text to form an image-text pair;
[0185] XI. Frame Extraction Service
[0186] (1) The download file system performs primary processing on the video file features of the file downloaded from the video content storage service - video frame extraction, including key frames and evenly extracted frames, as the input source for the cover image frames of the subsequent construction of the multi-modal pre-training model and the input for video OCR text recognition;
[0187] XII. Image Multi-modal Pre-training Database
[0188] Save the pre-training data corpus of the corresponding images crawled from the Internet, mainly by using Query words to retrieve public image data through a search engine;
[0189] Save the video frame data extracted from the text and image cover images, video cover images, and video content of the main information flow distribution channels, or the public data sets in the corresponding technical fields as the pre-training data corpus;
[0190] XIII. Crawling and Data Preprocessing System
[0191] Crawl the corresponding image data from the Internet according to the keywords constructed by the tags discovered through the information flow content described above;
[0192] Images corresponding to similar or identical tags can be considered similar.
[0193] This application can effectively model the recognition of multiple types of inappropriate pictures with a small amount of business sample data, improve the response and processing speed of inappropriate picture content problems, reduce the R & D cost of the model, and directly improve the cover picture experience of the object; by introducing the text modality, the context scenario of the picture can be fully considered, and the recognition effect is greatly improved; at the same time, it can provide support for the single picture modality, and good recognition effects can also be obtained for single pictures lacking text input, increasing the adaptability of the model, and good recognition can also be performed on the cover of small video content without titles.
[0194] Please refer to Figure 13 , this embodiment further provides an information classification device, which may include:
[0195] The first acquisition module 1310 is used to acquire the target information of the target image; the target information is the image information and text information of the target image, or the target information is the image information of the target image;
[0196] The first classification module 1320 is used to classify the target information based on the information classification model to obtain the first classification result of the target image; the information classification model is obtained by training the classification model to be trained based on the target graphic and text samples corresponding to the target business; the classification model to be trained is obtained based on the target feature extraction model and the preset classification layer; the target feature extraction model is obtained by respectively training multiple pre-trained task models based on the preset graphic and text samples, and the target feature extraction model is a shared model of the multiple pre-trained task models.
[0197] Furthermore, the preset graphic and text samples include multiple preset graphic and text matching items, and the image information and text information in each preset graphic and text matching item match each other;
[0198] The device further includes:
[0199] The information extraction module is used to respectively extract the information from the text information in the multiple preset graphic and text matching items to obtain the named entity information corresponding to the multiple preset graphic and text matching items respectively;
[0200] The similar graphic and text matching item determination module is used to determine the similar graphic and text matching items of each preset graphic and text matching item based on the named entity information corresponding to each preset graphic and text matching item and the named entity information corresponding to the remaining preset graphic and text matching items; the remaining preset graphic and text matching items are the preset graphic and text matching items in the multiple preset graphic and text matching items except each preset graphic and text matching item;
[0201] The first new matching item determination module is used to generate the first new matching item based on the image information in each preset graphic and text matching item and the text information in the similar graphic and text matching items;
[0202] A second newly added matching item determination module, configured to generate a second newly added matching item based on the image information in the similar graphic and text matching item and the text information in each preset graphic and text matching item.
[0203] Further, the multiple pre-training task models include a first task model, and the first task model includes an initial feature extraction model and a contrast output layer;
[0204] The apparatus further includes:
[0205] A first feature extraction module, configured to respectively perform feature extraction on the preset graphic and text sample and the corresponding contrast graphic and text sample based on the initial feature extraction model to obtain corresponding first output features and second output features; the contrast graphic and text sample is obtained by performing information transformation on the image information and / or text information in the preset graphic and text sample;
[0206] A contrast processing module, configured to perform contrast processing on the first output feature and the second output feature based on the contrast output layer to obtain contrast output information;
[0207] A first adjustment module, configured to adjust the model parameters of the first task model based on the contrast output information and the preset matching information between the preset graphic and text sample and the contrast graphic and text sample to obtain a trained first task model;
[0208] A first determination module, configured to determine the target feature extraction model based on the trained first task model.
[0209] Further, the multiple pre-training task models include a second task model, the preset graphic and text sample includes multiple preset graphic and text matching items, and the image information and text information in each preset graphic and text matching item are matched;
[0210] The apparatus further includes:
[0211] An information replacement module, configured to obtain a replacement graphic and text matching item corresponding to each of the multiple preset graphic and text matching items; the replacement graphic and text matching item is obtained by performing information replacement on the text information in the multiple preset graphic and text matching items;
[0212] A division module, configured to perform matching item division based on the multiple preset graphic and text matching items and the replacement graphic and text matching items corresponding to the multiple preset graphic and text matching items respectively to obtain a first replacement item and a second replacement item; the first replacement item includes replacement items with matching image information and text information and replacement items with non-matching image information and text information; the second replacement item is a replacement item with non-matching image information and text information;
[0213] The first training module is used to train the second task model based on the first replacement item to obtain a trained second task model;
[0214] The matching prediction module is used to perform matching prediction on the second replacement item based on the trained second task model to obtain matching prediction information;
[0215] The target negative sample determination module is used to determine the second replacement item as the target negative sample when the matching prediction information indicates that the image information and text information in the second replacement item match;
[0216] The second training module is used to train the trained second task model based on the target negative sample to obtain an updated second task model;
[0217] The second determination module is used to determine the target feature extraction model based on the updated second task model.
[0218] Furthermore, the second task model includes an initial feature extraction model and a matching output layer;
[0219] The second training module includes:
[0220] The second feature extraction module is used to extract features from the first replacement item based on the initial feature extraction model to obtain a third output feature;
[0221] The matching processing module is used to perform matching processing on the third output feature based on the matching output layer to obtain matching output information;
[0222] The second adjustment module is used to adjust the model parameters of the second task model based on the matching output information and the preset matching information of the image information and text information in the first replacement item to obtain the trained second task model.
[0223] Furthermore, the multiple pre-trained task models include a third task model, and the third task model includes an initial feature extraction model and a classification output layer; the preset graphic and text samples include multiple preset graphic and text matching items, and the image information and text information in each preset graphic and text matching item match;
[0224] The device further includes:
[0225] The third feature extraction module is used to extract features from the null information items corresponding to the multiple preset graphic and text matching items based on the initial feature extraction model; the null information items are obtained by performing null processing on the text information in the multiple preset graphic and text matching items;
[0226] A classification processing module, configured to perform classification processing on the fourth output feature based on the classification output layer to obtain classification output information;
[0227] A third adjustment module, configured to adjust the model parameters of the third task model based on the classification output information and the target classification labels of the multiple preset graphic-text matching items to obtain a trained third task model;
[0228] A third determination module, configured to determine the target feature extraction model based on the trained third task model.
[0229] Further, the target graphic-text sample includes multiple target classification labels;
[0230] The apparatus further includes:
[0231] A second acquisition module, configured to acquire an updated graphic-text sample; the updated classification label in the updated graphic-text sample is different from the multiple target classification labels;
[0232] A third training module, configured to perform model training on the information classification model based on the updated graphic-text sample to obtain an updated information classification model.
[0233] Please refer to Figure 14 , this embodiment further provides a multimedia resource pushing apparatus, which may include:
[0234] A candidate resource information determination module 1410, configured to determine candidate resource information corresponding to a candidate multimedia resource; the candidate resource information is image information and text information corresponding to the candidate multimedia resource, or the candidate resource information is image information corresponding to the multimedia resource;
[0235] A second classification module 1420, configured to classify the candidate resource information based on an information classification model to obtain a second classification result of the candidate multimedia resource; the information classification model is obtained by training a to-be-trained classification model based on a target graphic-text sample corresponding to a target service; the to-be-trained classification model is obtained based on a target feature extraction model and a preset classification layer; the target feature extraction model is obtained by performing model training on multiple pre-trained task models respectively based on a preset graphic-text sample, and the target feature extraction model is a shared model of the multiple pre-trained task models;
[0236] A target multimedia resource determination module 1430, configured to determine a target multimedia resource from the candidate multimedia resources based on the second classification result;
[0237] A resource pushing module 1440, configured to push the target multimedia resource.
[0238] Further, the second classification module 1420 includes:
[0239] A nulling processing module, configured to perform nulling processing on the text information corresponding to the candidate multimedia resource to obtain nulling information when the candidate resource information is the image information corresponding to the candidate multimedia resource;
[0240] An information classification result determination module, configured to input the image information corresponding to the candidate multimedia resource and the nulling information into the information classification model to obtain an information classification result of the candidate resource information; the information classification model is obtained by replacing or nulling the text information in the preset text and image samples.
[0241] The device provided in the above embodiment can execute the method provided in any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the above embodiment can be found in the method provided in any embodiment of the present application.
[0242] This embodiment also provides a computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to execute any of the above methods in this embodiment.
[0243] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes any of the above methods.
[0244] Figure 15 is a block diagram of an electronic device for an information classification method or a multimedia resource push method shown according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as Figure 15 shown. The electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an information classification method or a multimedia resource push method.
[0245] Those skilled in the art can understand, Figure 15The structure shown is only a block diagram of some of the structures related to the present disclosure, and does not constitute a limitation on the electronic device to which the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0246] This specification provides method operation steps as described in the embodiments or flowcharts, but may include more or fewer operation steps based on routine or non-creative labor. The steps and sequences listed in the embodiments are only one way among many execution sequences of the steps, and do not represent the only execution sequence. When the actual system or interrupt product is executed, it can be executed in the method sequence shown in the embodiments or the drawings or in parallel (for example, in an environment of parallel processors or multi-threaded processing).
[0247] The structure shown in this embodiment is only some of the structures related to the present application solution, and does not constitute a limitation on the device to which the present application solution is applied. The specific device may include more or fewer components than those shown, or combine certain components, or have a different component arrangement. It should be understood that the methods, devices, etc. disclosed in this embodiment can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, indirect coupling or communication connection of the device or unit module.
[0248] Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0249] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this specification can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0250] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application.
Claims
1. An information classification method, characterized in that, it includes: Obtaining target information of a target image; The target information is the image information and text information of the target image, or the target information is the image information of the target image; Classifying the target information based on an information classification model to obtain a first classification result of the target image; The information classification model is obtained by training a classification model to be trained based on target graphic and text samples corresponding to a target service; The classification model to be trained is obtained based on a target feature extraction model and a preset classification layer; the target feature extraction model is obtained by respectively training multiple pre-trained task models based on preset graphic and text samples, and the target feature extraction model is a shared model of the multiple pre-trained task models; the multiple pre-trained task models include the same feature extraction model, and the multiple pre-trained task models share the model parameters of the feature extraction model, and the training processes of the multiple pre-trained task models are various training constraints on the target feature extraction model; The training processes of the multiple pre-trained task models are carried out synchronously. During the synchronous training process, the continuously updated model parameters of the feature extraction model are shared among multiple pre-trained tasks.
2. The method according to claim 1, characterized in that, The preset graphic and text samples include multiple preset graphic and text matching items, and the image information and text information in each preset graphic and text matching item match; The method further includes: Respectively extracting information from the text information in the multiple preset graphic and text matching items to obtain named entity information corresponding to the multiple preset graphic and text matching items respectively; Based on the named entity information corresponding to each preset graphic and text matching item and the named entity information corresponding to the remaining preset graphic and text matching items, determining the similar graphic and text matching items of each preset graphic and text matching item; the remaining preset graphic and text matching items are the preset graphic and text matching items other than each preset graphic and text matching item among the multiple preset graphic and text matching items; Generating a first new matching item based on the image information in each preset graphic and text matching item and the text information in the similar graphic and text matching items; Generating a second new matching item based on the image information in the similar graphic and text matching items and the text information in each preset graphic and text matching item.
3. The method according to claim 1, characterized in that, The multiple pre-trained task models include a first task model, and the first task model includes an initial feature extraction model and a comparison output layer; The method further includes: Based on the initial feature extraction model, respectively extracting features from the preset graphic and text samples and the corresponding comparison graphic and text samples to obtain corresponding first output features and second output features; the comparison graphic and text samples are obtained by performing information transformation on the image information and / or text information in the preset graphic and text samples; Based on the comparison output layer, performing comparison processing on the first output feature and the second output feature to obtain comparison output information; Based on the comparison output information and the preset matching information between the preset graphic and text sample and the comparison graphic and text sample, adjust the model parameters of the first task model to obtain the trained first task model; Determine the target feature extraction model based on the trained first task model.
4. The method according to claim 1, wherein, the multiple pre-trained task models include a second task model, the preset graphic and text sample includes multiple preset graphic and text matching items, and the image information and text information in each preset graphic and text matching item match; the method further includes: Obtain the replacement graphic and text matching item corresponding to each of the multiple preset graphic and text matching items; the replacement graphic and text matching item is obtained by performing information replacement on the text information in the multiple preset graphic and text matching items; Based on the multiple preset graphic and text matching items and the replacement graphic and text matching items corresponding to each of the multiple preset graphic and text matching items, perform matching item division to obtain a first replacement item and a second replacement item; the first replacement item includes replacement items where the image information and text information match and replacement items where the image information and text information do not match; the second replacement item is a replacement item where the image information and text information do not match; Based on the first replacement item, perform model training on the second task model to obtain the trained second task model; Based on the trained second task model, perform matching prediction on the second replacement item to obtain matching prediction information; When the matching prediction information indicates that the image information and text information in the second replacement item match, determine the second replacement item as the target negative sample; Based on the target negative sample, perform model training on the trained second task model to obtain the updated second task model; Based on the updated second task model, determine the target feature extraction model.
5. The method according to claim 4, wherein, the second task model includes an initial feature extraction model and a matching output layer; The performing model training on the second task model based on the first replacement item to obtain the trained second task model includes: Based on the initial feature extraction model, perform feature extraction on the first replacement item to obtain a third output feature; Based on the matching output layer, perform matching processing on the third output feature to obtain matching output information; Based on the matching output information and the preset matching information between the image information and text information in the first replacement item, adjust the model parameters of the second task model to obtain the trained second task model.
6. The method according to claim 1, wherein, the multiple pre-trained task models include a third task model, the third task model includes an initial feature extraction model and a classification output layer; the preset graphic and text sample includes multiple preset graphic and text matching items, and the image information and text information in each preset graphic and text matching item match; the method further includes: Performing feature extraction on the null information items corresponding to the multiple preset text-image matching items based on the initial feature extraction model to obtain a fourth output feature; the null information items are obtained by performing nulling processing on the text information in the multiple preset text-image matching items. Performing classification processing on the fourth output feature based on the classification output layer to obtain classification output information. Adjusting the model parameters of the third task model based on the classification output information and the target classification labels of the multiple preset text-image matching items to obtain a trained third task model. Determining the target feature extraction model based on the trained third task model.
7. The method according to claim 1, wherein, the target text-image sample includes multiple target classification labels; the method further includes: Obtaining an updated text-image sample; the updated classification label in the updated text-image sample is different from the multiple target classification labels. Performing model training on the information classification model based on the updated text-image sample to obtain an updated information classification model.
8. A method for pushing multimedia resources, wherein, it includes: Determining candidate resource information corresponding to a candidate multimedia resource; the candidate resource information is the image information and text information corresponding to the candidate multimedia resource, or the candidate resource information is the image information corresponding to the multimedia resource; Classifying the candidate resource information based on an information classification model to obtain a second classification result of the candidate multimedia resource; the information classification model is obtained by training a classification model to be trained based on a target text-image sample corresponding to a target service; the classification model to be trained is obtained based on a target feature extraction model and a preset classification layer; the target feature extraction model is obtained by performing model training on multiple pre-training task models respectively based on a preset text-image sample, and the target feature extraction model is a shared model of the multiple pre-training task models; the multiple pre-training task models include the same feature extraction model, and the multiple pre-training task models share the model parameters of the feature extraction model, and the training processes of the multiple pre-training task models are multiple training constraints on the target feature extraction model; the training processes of the multiple pre-training task models are performed synchronously, and during the synchronous training process, the model parameters of the continuously updated feature extraction model are shared among the multiple pre-training tasks. Determining a target multimedia resource from the candidate multimedia resources based on the second classification result. Pushing the target multimedia resource.
9. The method according to claim 8, wherein, the classifying the candidate resource information based on the information classification model to obtain the second classification result of the candidate multimedia resource includes: In the case where the candidate resource information is the image information corresponding to the candidate multimedia resource, performing nulling processing on the text information corresponding to the candidate multimedia resource to obtain null information. Input the image information corresponding to the candidate multimedia resource and the nulling information into the information classification model to obtain the second classification result; the information classification model is obtained by replacing or nulling the text information in the preset text-image sample.
10. An information classification device, characterized in that, it includes: A first acquisition module, configured to acquire the target information of a target image; The target information is the image information and text information of the target image, or the target information is the image information of the target image; A first classification module, configured to classify the target information based on an information classification model to obtain a first classification result of the target image; The information classification model is obtained by training a classification model to be trained based on a target text-image sample corresponding to a target service; The classification model to be trained is obtained based on a target feature extraction model and a preset classification layer; the target feature extraction model is obtained by training multiple pre-trained task models respectively based on a preset text-image sample, and the target feature extraction model is a shared model of the multiple pre-trained task models; the multiple pre-trained task models include the same feature extraction model, and the multiple pre-trained task models share the model parameters of the feature extraction model, and the training processes of the multiple pre-trained task models are various training constraints on the target feature extraction model; The training processes of the multiple pre-trained task models are carried out synchronously. During the synchronous training process, the continuously updated model parameters of the feature extraction model are shared among multiple pre-trained tasks.
11. The device according to claim 10, characterized in that, the preset text-image sample includes multiple preset text-image matching items, and the image information and text information in each preset text-image matching item match; the device further includes: An information extraction module, configured to respectively extract information from the text information in the multiple preset text-image matching items to obtain named entity information corresponding to the multiple preset text-image matching items respectively; A similar text-image matching item determination module, configured to determine a similar text-image matching item of each preset text-image matching item based on the named entity information corresponding to each preset text-image matching item and the named entity information corresponding to the remaining preset text-image matching items; the remaining preset text-image matching items are the preset text-image matching items other than each preset text-image matching item in the multiple preset text-image matching items; A first new matching item determination module, configured to generate a first new matching item based on the image information in each preset text-image matching item and the text information in the similar text-image matching item; A second new matching item determination module, configured to generate a second new matching item based on the image information in the similar text-image matching item and the text information in each preset text-image matching item.
12. The device according to claim 10, characterized in that, the multiple pre-trained task models include a first task model, and the first task model includes an initial feature extraction model and a comparison output layer; the device further includes: The first feature extraction module is configured to perform feature extraction on the preset graphic and text sample and the corresponding comparison graphic and text sample based on the initial feature extraction model, so as to obtain corresponding first output features and second output features; the comparison graphic and text sample is obtained by performing information transformation on the image information and / or text information in the preset graphic and text sample. The comparison processing module is configured to perform comparison processing on the first output features and the second output features based on the comparison output layer to obtain comparison output information. The first adjustment module is configured to adjust the model parameters of the first task model based on the comparison output information and the preset matching information between the preset graphic and text sample and the comparison graphic and text sample, so as to obtain a trained first task model. The first determination module is configured to determine the target feature extraction model based on the trained first task model.
13. The apparatus according to claim 10, wherein, the multiple pre-trained task models include a second task model, the preset graphic and text sample includes multiple preset graphic and text matching items, and the image information and text information in each preset graphic and text matching item match; the apparatus further includes: The information replacement module is configured to obtain respective replacement graphic and text matching items corresponding to the multiple preset graphic and text matching items; the replacement graphic and text matching items are obtained by performing information replacement on the text information in the multiple preset graphic and text matching items. The division module is configured to perform matching item division based on the multiple preset graphic and text matching items and the respective replacement graphic and text matching items corresponding to the multiple preset graphic and text matching items, so as to obtain a first replacement item and a second replacement item; the first replacement item includes replacement items in which the image information and text information match and replacement items in which the image information and text information do not match; the second replacement item is a replacement item in which the image information and text information do not match. The first training module is configured to perform model training on the second task model based on the first replacement item to obtain a trained second task model. The matching prediction module is configured to perform matching prediction on the second replacement item based on the trained second task model to obtain matching prediction information. The target negative sample determination module is configured to determine the second replacement item as a target negative sample when the matching prediction information indicates that the image information and text information in the second replacement item match. The second training module is configured to perform model training on the trained second task model based on the target negative sample to obtain an updated second task model. The second determination module is configured to determine the target feature extraction model based on the updated second task model.
14. The apparatus according to claim 13, wherein, the second task model includes an initial feature extraction model and a matching output layer; the second training module includes: The second feature extraction module is configured to perform feature extraction on the first replacement item based on the initial feature extraction model to obtain third output features. The matching processing module is configured to perform matching processing on the third output features based on the matching output layer to obtain matching output information. A second adjustment module, configured to adjust model parameters of the second task model based on the matching output information and preset matching information between the image information and the text information in the first replacement item, so as to obtain the trained second task model.
15. The apparatus according to claim 10, wherein, the multiple pre-trained task models include a third task model, and the third task model includes an initial feature extraction model and a classification output layer; the preset graphic and text samples include multiple preset graphic and text matching items, and the image information and the text information in each preset graphic and text matching item match each other; the apparatus further includes: a third feature extraction module, configured to extract features from the null information items corresponding to the multiple preset graphic and text matching items based on the initial feature extraction model, so as to obtain a fourth output feature; the null information items are obtained by performing a null processing on the text information in the multiple preset graphic and text matching items; a classification processing module, configured to perform classification processing on the fourth output feature based on the classification output layer, so as to obtain classification output information; a third adjustment module, configured to adjust model parameters of the third task model based on the classification output information and target classification labels of the multiple preset graphic and text matching items, so as to obtain a trained third task model; a third determination module, configured to determine the target feature extraction model based on the trained third task model.
16. The apparatus according to claim 10, wherein, the target graphic and text sample includes multiple types of target classification labels; the apparatus further includes: a second acquisition module, configured to acquire an updated graphic and text sample; the updated classification label in the updated graphic and text sample is different from the multiple types of target classification labels; a third training module, configured to perform model training on the information classification model based on the updated graphic and text sample, so as to obtain an updated information classification model.
17. A multimedia resource push apparatus, wherein, it includes: a candidate resource information determination module, configured to determine candidate resource information corresponding to a candidate multimedia resource; the candidate resource information is image information and text information corresponding to the candidate multimedia resource, or the candidate resource information is image information corresponding to the multimedia resource; a second classification module, configured to classify the candidate resource information based on an information classification model, so as to obtain a second classification result of the candidate multimedia resource; the information classification model is obtained by training a to-be-trained classification model based on a target graphic and text sample corresponding to a target service; the to-be-trained classification model is obtained based on a target feature extraction model and a preset classification layer; the target feature extraction model is obtained by performing model training on multiple pre-trained task models respectively based on a preset graphic and text sample, and the target feature extraction model is a shared model of the multiple pre-trained task models; the multiple pre-trained task models include the same feature extraction model, and the multiple pre-trained task models share model parameters of the feature extraction model, and the training processes of the multiple pre-trained task models are multiple training constraints on the target feature extraction model. The training processes of the multiple pre-trained task models are carried out synchronously. During the synchronous training process, the model parameters of the continuously updated feature extraction model are shared among multiple pre-trained tasks; A target multimedia resource determination module, configured to determine a target multimedia resource from the candidate multimedia resources based on the second classification result; A resource push module, configured to push the target multimedia resource.
18. The apparatus according to claim 17, wherein, the second classification module includes: A nulling processing module, configured to perform nulling processing on the text information corresponding to the candidate multimedia resource to obtain nulling information when the candidate resource information is the image information corresponding to the candidate multimedia resource; An information classification result determination module, configured to input the image information corresponding to the candidate multimedia resource and the nulling information into the information classification model to obtain an information classification result of the candidate resource information; the information classification model is obtained by replacing or nulling the text information in the preset text-image sample.
19. An electronic device, wherein, the device includes a processor and a memory, and at least one instruction or at least one segment of program is stored in the memory. The at least one instruction or the at least one segment of program is loaded and executed by the processor to implement the information classification method according to any one of claims 1 to 7, or the multimedia resource push method according to any one of claims 8 to 9.
20. A computer storage medium, wherein, at least one instruction or at least one segment of program is stored in the storage medium. The at least one instruction or the at least one segment of program is loaded and executed by a processor to implement the information classification method according to any one of claims 1 to 7, or the multimedia resource push method according to any one of claims 8 to 9.
21. A computer program product or a computer program, wherein, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the information classification method according to any one of claims 1 to 7, or the multimedia resource push method according to any one of claims 8 to 9.
Citation Information
Patent Citations
Media information classification method, method and apparatus for training picture classification model
CN109344884A
Video classification method, device and equipment based on multi-modal representation, and storage medium
CN113762322A