Content resource detection model training method, content resource detection method and device

By training the basic models of the target field and combining the sample labeling data of the target business and resource quality detection business, using self-supervised and weak supervision comparison learning methods, the problem of poor generalization ability of the content resource detection model is solved, and efficient and accurate content resource detection is achieved.

CN116933069BActive Publication Date: 2025-08-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210346752.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-08-08
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

Due to the limited sample content, the existing content resource detection model has poor generalization capabilities and low detection accuracy.

Method used

By training the basic models in the target field, combining the supervision methods of the target business and the sample labeling data of the resource quality inspection services, the model is refined layer by layer, and the self-supervision and weak supervision comparison learning methods are used to improve the model performance.

Benefits of technology

It improves the accuracy and generalization of content resource detection, shortens model development time, and reduces manual labeling costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116933069B_ABST
    Figure CN116933069B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for a content resource detection model, a content resource detection method, a system, an apparatus, a computer device and a storage medium, and belongs to the field of artificial intelligence technology. This application is based on the data of the target field, trains a basic model that supports multiple CV tasks in the target field, and then uses the data of the target business to train the basic model for pertinence in the business scenario, thereby obtaining a first detection model. While ensuring the performance of the model, the generalization of the first detection model is effectively improved, thereby using a small amount of labeled data corresponding to the target resource quality detection task to train the first detection model, and the content resource detection model corresponding to the target resource quality detection task can be quickly obtained. From a large range of fields to business scenarios and then to downstream specific businesses, the training method of the model is refined layer by layer, and the different samples corresponding to each level are fully utilized during the model training process, which greatly improves the accuracy of content resource detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a content resource detection model training method, content resource detection method, system, device, computer equipment and storage medium. Background Art

[0002] Graphics and text content is a common form of content dissemination on the internet. Graphics and text content consists of both images and text. Before distributing content, it's typically inspected to filter out low-quality content, such as unclear or incomplete images or content that doesn't conform to the text's semantics.

[0003] Due to the huge amount of data, manual detection is inefficient. Currently, the pre-trained model is usually fine-tuned based on the large amount of labeled sample content corresponding to the detection business to obtain a content resource detection model to assist in detection.

[0004] However, since there are many types of content resource-related detection services, and the sample content that each service type can provide is very limited, the content resource detection model obtained in the above technical solution will overfit to a small amount of sample content, and the model's generalization ability is very poor, resulting in very low accuracy in content resource detection. Summary of the Invention

[0005] The embodiments of the present application provide a content resource detection model training method, content resource detection method, system, apparatus, computer equipment, and storage medium, which effectively improve the accuracy of content resource detection. The technical solution is as follows:

[0006] In one aspect, a method for training a content resource detection model is provided, the method comprising:

[0007] Training a basic model for the target domain based on a first sample corresponding to the target domain, the first sample including content resources associated with keywords in the target domain, the basic model including multiple computer vision tasks associated with the target domain;

[0008] Based on a second sample corresponding to a target business in the target domain, the basic model is trained using a supervision method corresponding to the second sample to obtain a first detection model corresponding to the target business, the target business corresponding to a target computer vision task, the second sample including an image sample and an image-text sample pair corresponding to the target business, the image-text sample pair including an image and text associated with the image;

[0009] Based on the third sample corresponding to the target resource quality detection business under the target business and the label of the third sample, the first detection model is trained to obtain a content resource detection model corresponding to the target resource quality detection business. The third sample includes content resources labeled based on the resource quality detection dimension corresponding to the target resource quality detection business.

[0010] In one aspect, a content resource detection method is provided, the method comprising:

[0011] Obtain the content resources to be tested in the target field;

[0012] Obtain the content resource detection model corresponding to the target resource quality detection business under the target business under the target field, and detect the content resource to be detected based on the content resource detection model. The target resource quality detection business corresponds to the resource quality detection dimension of the content resource, and the content resource detection model is obtained based on the training method of the above-mentioned content resource detection model.

[0013] In one aspect, a training device for a content resource detection model is provided, the device comprising:

[0014] A first training module is configured to train a basic model for a target domain based on a first sample corresponding to the target domain, the first sample including content resources associated with keywords in the target domain, and the basic model including multiple computer vision tasks associated with the target domain;

[0015] a second training module, configured to train the basic model based on a second sample corresponding to a target business in the target domain, using a supervision method corresponding to the second sample, to obtain a first detection model corresponding to the target business, wherein the target business corresponds to a target computer vision task, and the second sample includes an image sample and an image-text sample pair corresponding to the target business, wherein the image-text sample pair includes an image and text associated with the image;

[0016] The third training module is used to train the first detection model based on the third sample corresponding to the target resource quality detection business under the target business and the label of the third sample, to obtain the content resource detection model corresponding to the target resource quality detection business. The third sample includes content resources labeled based on the resource quality detection dimension corresponding to the target resource quality detection business.

[0017] In one possible implementation, the second training module includes:

[0018] A first training unit is configured to, when the second sample is an image sample, train the basic model using a self-supervised contrastive learning method based on the second sample to obtain the first detection model;

[0019] The second training unit is used to train the basic model by adopting a weakly supervised contrast learning method based on the second sample to obtain the first detection model when the second sample is a picture-text sample pair.

[0020] In one possible implementation, the second training module includes:

[0021] A loss function determining unit, configured to determine a target loss function corresponding to the target computer vision task based on the target business;

[0022] The first training unit is configured to, when the second sample is an image sample, train the basic model using the self-supervised contrastive learning method based on the second sample, the target computer vision task, and the target loss function to obtain the first detection model;

[0023] The second training unit is used to train the basic model using the weakly supervised contrast learning method based on the second sample, the target computer vision task and the target loss function to obtain the first detection model when the second sample is a picture-text sample pair.

[0024] In one possible implementation, the first training unit is used to:

[0025] Input the image sample into the basic model to perform the target computer vision task;

[0026] Based on the first output result of the target computer vision task, using a self-supervised contrastive learning method, calculating the target loss function and determining a first loss value, where the first loss value represents an error in learning the features of the image sample by the first detection model;

[0027] Based on the first loss value, the first detection model is trained.

[0028] In one possible implementation, the second training unit:

[0029] Input the image-text sample pair into the basic model to perform the target computer vision task;

[0030] Based on the second output result of the target computer vision task, using a weakly supervised contrastive learning method, calculating the target loss function to obtain a second loss value, where the second loss value represents an error in the first detection model's prediction of the similarity between the image and the text in the image-text sample pair;

[0031] The first detection model is trained based on the second loss value.

[0032] In one possible implementation, if the target computer vision task is a classification task or a recognition task, the target loss function is a reconstruction loss function or a simulation loss function;

[0033] The reconstruction loss function is used to enable the first detection model to learn the ability to predict the missing part of the image; the simulation loss function is used to enable the first detection model to learn the parameters of the base model;

[0034] If the computer vision task is a segmentation task or a detection task, the target loss function is a dense contrast loss function;

[0035] The dense contrast loss function is used to enable the first detection model to learn pixel-level features of the image.

[0036] In one possible implementation, the second sample includes:

[0037] A picture sample determined based on a frame-drawing picture obtained from a content resource of the target service, the picture sample including a positive picture sample constructed based on a plurality of similar frame-drawing pictures, and a negative picture sample constructed based on a plurality of dissimilar frame-drawing pictures;

[0038] The image-text sample pairs are determined based on the images and texts included in the content resources of the target business.

[0039] In one possible implementation, the basic model of the target domain is constructed based on a general backbone model, which includes but is not limited to any of the following:

[0040] Deep residual model, visual transformation model, pyramid visual transformation model, convolutional biased visual transformation model or sliding window layered visual transformation model.

[0041] In one aspect, a content resource detection device is provided, the device comprising:

[0042] The acquisition module is used to obtain the content resources to be detected in the target field;

[0043] The detection module is used to obtain the content resource detection model corresponding to the target resource quality detection business under the target business under the target field, and detect the content resource to be detected based on the content resource detection model. The target resource quality detection business corresponds to the resource quality detection dimension of the content resource, and the content resource detection model is obtained based on the training method of the above-mentioned content resource detection model.

[0044] In one aspect, a content resource detection system is provided, the system comprising:

[0045] A first database is configured to store a first sample corresponding to a target domain and a second sample corresponding to a target business within the target domain, wherein the first sample includes content resources associated with keywords of the target domain, and the second sample includes an image sample and an image-text sample pair corresponding to the target business, wherein the image-text sample pair includes an image and text associated with the image;

[0046] A second database is used to store a third sample corresponding to a target resource quality detection service under the target service and a label of the third sample, where the third sample includes content resources annotated based on a resource quality detection dimension corresponding to the target resource quality detection service;

[0047] The server is configured to train a basic model for the target domain based on the first sample, the basic model including multiple computer vision tasks associated with the target domain; train the basic model based on a second sample corresponding to a target business under the target domain using a supervision method corresponding to the second sample to obtain a first detection model corresponding to the target business, the target business corresponding to the target computer vision task; and train the first detection model based on a third sample corresponding to a target resource quality detection business under the target business and a label of the third sample to obtain a content resource detection model corresponding to the target resource quality detection business;

[0048] The server is used to obtain the content resources to be detected in the target field;

[0049] A third database is used to store the content resource to be detected and metadata of the content resource to be detected, where the metadata includes relevant information of the content resource to be detected;

[0050] The server is used to schedule the content resource detection model to detect the content resource to be detected based on the third database.

[0051] On the one hand, a computer device is provided, which includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the above-mentioned content resource detection model training method or content resource detection method.

[0052] On the one hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned content resource detection model training method or content resource detection method.

[0053] On the one hand, a computer program product or computer program is provided, which includes a program code, which is stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above-mentioned content resource detection model training method or content resource detection method.

[0054] This application uses data from the target domain as a basis to train a basic model that supports multiple CV tasks in the target domain. The target business data is then used to train the basic model for pertinence in business scenarios, resulting in a first detection model. This effectively improves the generalization of the first detection model while ensuring model performance. By using a small amount of labeled data corresponding to the target resource quality detection task to train the first detection model, the content resource detection model corresponding to the target resource quality detection task can be quickly obtained. The model training method is refined layer by layer, from large-scale domains to business scenarios to downstream specific businesses. The model training process fully utilizes different samples corresponding to each level, greatly improving the accuracy of content resource detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0056] Figure 1 is a schematic diagram of a content resource detection system provided in an embodiment of the present application;

[0057] Figure 2 This is a flowchart of a method for training a content resource detection model provided in an embodiment of the present application;

[0058] Figure 3 This is a flowchart of a method for training a content resource detection model provided in an embodiment of the present application;

[0059] Figure 4 is a schematic diagram of dense contrastive learning provided in an embodiment of the present application;

[0060] Figure 5 Schematic diagram of a self-supervised contrast loss function provided in an embodiment of the present application;

[0061] Figure 6 is a schematic diagram of training a first detection model provided in an embodiment of the present application;

[0062] Figure 7This is a schematic diagram of processing a graphic sample pair provided by an embodiment of the present application;

[0063] Figure 8 is a schematic diagram of a training method for a content resource detection model provided in an embodiment of the present application;

[0064] Figure 9 is a schematic diagram of a training method for a content resource detection model provided in an embodiment of the present application;

[0065] Figure 10 This is a flow chart of a content resource detection method provided by an embodiment of the present application;

[0066] Figure 11 This is a schematic diagram of a content resource distribution process provided by an embodiment of the present application;

[0067] Figure 12 This is a structural diagram of a training device for a content resource detection model provided in an embodiment of the present application;

[0068] Figure 13 This is a structural diagram of a content resource detection device provided in an embodiment of the present application;

[0069] Figure 14 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0070] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0071] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity and execution order.

[0072] In the present application, the term "at least one" means one or more, and the term "plurality" means two or more. For example, a plurality of pictures means two or more pictures.

[0073] The technical solution provided in this application involves the field of artificial intelligence and can be applied to various scenarios such as image processing, cloud technology, and big data.

[0074] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning.

[0075] Computer vision (CV) technology is the study of how machines can "see." Specifically, it refers to machine vision, such as using cameras and computers to replace the human eye to identify and measure objects. Further image processing is performed to transform the computer-generated images into images more suitable for human observation or transmission to instrumentation. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, and simultaneous localization and mapping.

[0076] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge sub-models to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.

[0077] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool for on-demand, flexible and convenient use. Cloud computing technology will become a key support. The backend services of technical network systems, such as video websites, image websites, and more portals, require extensive computing and storage resources. With the rapid development and application of the internet industry, every item will likely have its own unique identification mark, which will need to be transmitted to the backend system for logical processing. Different levels of data will be processed separately. All kinds of industry data require strong system support, which can only be achieved through cloud computing.

[0078] A database, in short, can be thought of as a digital filing cabinet—a place where electronic files are stored, allowing users to add, query, update, and delete data. A database is a collection of data stored in a specific way, shared by multiple users, with minimal redundancy, and independent of applications.

[0079] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the content resources involved in this application were obtained with full authorization.

[0080] Figure 1 This is a schematic diagram of a content resource detection system provided by an embodiment of the present application, see Figure 1 The content resource detection system 100 includes a server 110 , a first database 120 , a second database 130 and a third database 140 .

[0081] Among them, the server 110 is used to obtain a first sample corresponding to the target field, and the first sample includes content resources associated with keywords of the target field. In some embodiments, the content resource is a multimedia resource, such as a short video, a picture collection, or graphic information; the target field refers to the dissemination field corresponding to the content resource, for example, the information flow content distribution field that disseminates content resources such as picture collections and short videos through the Internet media. Correspondingly, the keywords of the target field refer to the keywords used to describe content resources in the target field, for example, keywords for describing pictures: scenery, animals, or plants; keywords for describing videos: beauty, cooking, or games. Based on this, the server can search according to the keywords of the target field to obtain the first sample.

[0082] Among them, the server 110 is also used to obtain a second sample corresponding to the target business in the target field, and the second sample includes a picture sample and a picture-text sample pair corresponding to the target business, and the picture-text sample pair includes a picture and a text associated with the picture. In some embodiments, the target business refers to the detection of content resources in the target content resource dissemination scenario in the target field, for example, in the scenario of short video content distribution, quality detection of the short video to be distributed. Accordingly, the content resources corresponding to the target business include detection objects targeted in the target scenario, for example, short videos and picture albums on short video platforms; videos and picture-text information on social platforms. In this example, the picture sample corresponding to the target business can be a picture in a picture album, and the picture-text sample pair included in the second sample can be the cover picture of the short video and the title text of the short video (the title is usually associated with the cover picture of the short video).

[0083] The first database 120 is used to store the first sample and the second sample. Optionally, the first database 120 also stores samples corresponding to other businesses in the target domain, so that the business scenarios covered by the first database 120 are richer.

[0084] Among them, the second database 130 is used to store the third sample corresponding to the target resource quality detection business under the target business and the label of the third sample. In some embodiments, the target resource quality detection business corresponds to the resource quality detection dimension of the content resource, and the resource quality detection dimension refers to a certain type of attribute of the content resource, such as picture clarity, picture completeness, picture-text consistency, or picture content tendency. Accordingly, the target resource quality detection business refers to the detection of the above-mentioned resource quality detection dimension on the content resource in the target content resource dissemination scenario corresponding to the target business, for example, in the scenario of short video distribution, the picture clarity of the short video to be distributed is detected.

[0085] Among them, the above-mentioned third sample includes content resources that are annotated based on the resource quality detection dimension corresponding to the target resource quality detection business. Accordingly, the label of the third sample indicates its detection result on the resource quality detection dimension. In some embodiments, the label of the third sample is a quality label corresponding to the resource quality detection dimension. For example, if the resource quality detection dimension is image integrity, the label of the third sample is a completeness label, and the completeness label indicates the degree of completeness of the content resource, such as complete, partially missing, or seriously missing; if the resource quality detection dimension is image semantic content tendency, the label of the third sample is a content tendency label, and the content tendency label indicates the content tendency of the content resource, such as visual stimulation tendency, unconventional tendency, or bad tendency.

[0086] In some embodiments, the server and staff determine the quality labels corresponding to the above-mentioned resource quality detection dimensions. For example, the content resources included in the third sample are short videos. The server detects the short videos based on the resource quality detection dimensions corresponding to the target resource quality detection business to obtain quality labels. The staff reviews the quality labels obtained by the server, and then the server returns the quality labels reviewed by the staff to the second database 130 for storage.

[0087] In other embodiments, the staff determines the quality label corresponding to the above-mentioned resource quality detection dimension. For example, in the scenario of short videos, the staff detects the quality of the short video targeted by the quality feedback based on the quality feedback of the short video, and determines the quality label of the short video. The server then returns the quality label of the short video determined by the staff to the second database 130 for storage.

[0088] Among them, the server 110 is used to obtain the first sample from the first database 120, and based on the first sample, train the basic model of the target domain, and the basic model includes multiple categories of computer vision tasks associated with the target domain; obtain the second sample from the first database 120; based on the second sample, use the supervision method corresponding to the second sample to train the basic model to obtain the first detection model corresponding to the target business, and the target business corresponds to the target computer vision task; obtain the third sample and the label of the third sample from the second database 130, and train the first detection model based on the third sample and the label of the third sample to obtain the content resource detection model corresponding to the target resource quality detection business.

[0089] Among them, the third database 140 is used to store the content resources to be detected and the metadata of the content resources to be detected, and the metadata includes relevant information of the content resources to be detected. In some embodiments, the metadata includes basic attribute information of the content resource to be detected, as well as basic content information of the content resource to be detected. For example, the content resource is a video, and the basic attribute information of the video includes: the size of the video file, the file format of the video, and the bit rate of the video; the basic content information of the video includes: the video cover image (which can be stored in the form of a link), the title of the video, and the creation object information of the video. In some embodiments, the metadata can be used to perform a preliminary classification of the content resource to be detected. For example, if the content resource to be detected is a performance analysis video of a smartphone, then based on the title of the video and the cover image of the video, the multi-level classification label of the video can be: technology - smartphone - mobile phone brand A - mobile phone model B.

[0090] Optionally, the above-mentioned process of preliminary classification of the content resources to be detected based on metadata can be performed by the above-mentioned server 110, or can be completed jointly by computer equipment and staff. For example, the server performs semantic splitting based on the video title to obtain rough classification information, and then the staff obtains detailed classification information based on the video cover image and the actual content of the video. This application does not limit this.

[0091] In some embodiments, the classification information obtained through the preliminary classification process, such as the multi-level classification labels of the video, is stored in the third database 140 as part of the metadata of the content resource. Optionally, the server 110 can dispatch the classification information to the third database 140 for storage, which is not limited in this embodiment of the present application.

[0092] In some embodiments, the server 110 is further used to obtain the content resources to be detected in the target field and schedule them to be stored in the third database 140. Optionally, the server 110 can read the metadata in the third database 140 and schedule the content resources to be detected to the server function module responsible for preliminary classification, or to the terminal where the staff responsible for preliminary classification is located. Optionally, the server 110 can determine the scheduling priority between the server function module and the terminal based on the actual situation of the content resources. For example, content resources released for the first time are scheduled to the server function module first; content resources that have been fed back with quality issues are scheduled to the terminal where the staff is located first. Optionally, the server function module can be a scheduling system responsible for scheduling each server node in the server 110.

[0093] The server 110 is configured to schedule the content resource detection model corresponding to the target resource quality detection dimension, and detect the content resource to be detected based on the third database 140 .

[0094] Optionally, the server 110 can obtain the above content resources from a terminal corresponding to a production object of the content resources.

[0095] It should be noted that the content resources in this application are obtained with full authorization. For example, when the producer of the short video selects the "Agree to Publish" option on the short video upload page, the server obtains the specified short video that the producer agrees to publish.

[0096] Optionally, the server 110 is further configured to pre-process the content resources, for example, selecting and taking screenshots of multiple images included in the atlas. In some embodiments, the server 110 can perform the above pre-processing based on metadata corresponding to the content resources, for example, extracting images frame by frame from the video and taking screenshots of the atlas based on the type of the content resources.

[0097] Optionally, the third database 110 is further used to store content resources to be distributed, and the content resources to be distributed and the content resources to be detected are stored in mutually isolated storage spaces. The content resources to be distributed can be content resources that meet quality requirements after being detected by the content resource detection model.

[0098] Optionally, when the server 110 stores the content resources in any of the above-mentioned databases, the server 110 can control the speed and progress of downloading the content resources. For example, the server performs distributed cache acceleration when downloading videos.

[0099] Optionally, server 110 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0100] Optionally, the first database 120 , the second database 130 , and the third database 140 may be independent physical servers, or a server cluster or distributed system composed of multiple physical servers, for providing data storage services.

[0101] In some embodiments, the terminal can be connected to the server 110 in the content resource detection system 100 via a wireless network or a wired network. Optionally, the terminal is installed with an application that supports uploading, downloading, or online browsing of content resources. Optionally, the terminal is a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to these.

[0102] In some embodiments, the server 110 is a node in the blockchain system. After the server 110 trains and obtains a content resource detection model, the server 110 publishes the content resource detection model to the blockchain system, i.e., stores it in the form of a block in the blockchain, thereby enabling other nodes in the blockchain system to apply the content resource detection model. Optionally, the first database 120, the second database 130, and the third database 140 can all serve as nodes in the blockchain system, used to store data such as content resources and detection results of content resources.

[0103] After introducing the implementation environment of the embodiment of the present application, the application scenario of the embodiment of the present application will be introduced in combination with the above-mentioned content resource detection system. In the following description, the server is also the above-mentioned server 110.

[0104] The technical solution provided by the embodiments of this application can be applied in content resource distribution scenarios to detect acquired content resources before distribution, thereby filtering out low-quality content resources. For example, in a short video platform, before a server aggregates multiple short video information streams and distributes them, the server performs quality detection based on the short video title text, cover image, and copy, filtering out low-quality short videos in the short video information stream.

[0105] After introducing a content resource detection system and application scenarios provided by this application, the training method of the content resource detection model provided by an embodiment of this application is described below. Figure 2 This is a flowchart of a method for training a content resource detection model provided in an embodiment of the present application, which is applied to the above-mentioned content resource detection system 100 and executed by the server 110. Figure 2 , methods include:

[0106] 201. The server trains a basic model of the target domain based on a first sample corresponding to the target domain, where the first sample includes content resources associated with keywords of the target domain, and the basic model includes multiple types of computer vision tasks associated with the target domain.

[0107] The definitions of the content resources, target domain and the first sample are as follows: Figure 1 The description of the corresponding content resource detection system is omitted here.

[0108] Computer Vision (CV) tasks refer to methods for acquiring, processing, analyzing, and understanding digital images in the process of understanding the content of an image or video. In some embodiments, the multiple CV tasks include detection tasks, segmentation tasks, classification tasks, and recognition tasks. In other embodiments, the CV tasks include tasks involving multiple data types, such as semantic alignment tasks for an image and text describing the image.

[0109] In some embodiments, in the process of detecting content resources corresponding to the target field, multiple types of CV tasks are performed to understand the content resources. For example, if the content resource is a picture, the target part (such as an object or a person) in the picture will be detected first, and then the target part will be segmented out, and the segmented target part will be classified to identify more specific information.

[0110] In an embodiment of the present application, the first sample corresponding to the target domain is used to train the basic model of the target domain. This is equivalent to training a general multi-category CV task involved in detecting content resources in the target domain with the support of data from the target domain, thereby ensuring that the basic model can support multiple types of CV tasks in parallel in the target domain, effectively improving the basic model's pertinence to the target domain, and thereby improving the accuracy of content resource detection in the target domain.

[0111] 202. The server trains the basic model based on a second sample corresponding to the target business in the target field and adopts a supervision method corresponding to the second sample to obtain a first detection model corresponding to the target business. The target business corresponds to a target computer vision task. The second sample includes an image sample and an image-text sample pair corresponding to the target business. The image-text sample pair includes an image and text associated with the image.

[0112] The definitions of the target service, the second sample, the image sample, and the image-text sample pair are as follows: Figure 1 The description of the corresponding content resource detection system is omitted here.

[0113] In an embodiment of the present application, different types of samples in the second sample correspond to different supervision methods. In some embodiments, the second sample is a picture sample, and the focus is on training the basic model to capture the features of the picture itself that are not affected by other factors, such as picture texture, contour, edge and structure. In this example, based on the picture sample, a self-supervised learning method is used to train the basic model, so that the basic model learns to extract the features of the picture. In other embodiments, the second sample is a picture-text sample pair, and the focus is on training the basic model to learn the semantic association between pictures and text, for example, whether the cover picture and title of a video match. In this example, based on the picture-text sample pair, a weakly supervised learning method is used to train the basic model, so that the basic model learns the alignment of pictures and text in the semantic space.

[0114] Among them, the target business refers to the detection of content resources in the target content resource dissemination scenario in the target field; the target computer vision task refers to: when detecting content resources in the target content resource dissemination scenario, the CV tasks involved, for example, in the scenario of short video content distribution, when performing quality inspection on the short videos to be distributed, it involves classification tasks and recognition tasks for images, then the target computer vision task includes the classification task and recognition task.

[0115] It can be understood that on the basis of the basic model including multiple categories of CV tasks, based on the target computer vision task, the second sample corresponding to the target business is used to train the basic model, which can further obtain a first detection model with good performance for the target business. While improving the accuracy of content resource detection, the first detection model also has good generalization within the scope of the target business.

[0116] 203. The server trains the first detection model based on the third sample corresponding to the target resource quality detection business under the target business and the label of the third sample to obtain a content resource detection model corresponding to the target resource quality detection business. The third sample includes content resources labeled based on the resource quality detection dimension corresponding to the target resource quality detection business.

[0117] The definitions of the target resource quality detection service, the third sample, and the resource quality detection dimension are as follows: Figure 1 The description of the corresponding content resource detection system is omitted here.

[0118] In the embodiments of the present application, the resource quality detection dimension refers to a certain type of attribute of the content resource. It can be understood that the target resource quality detection service refers to: in the target content resource dissemination scenario corresponding to the target service, detecting a certain type of attribute of the content resource to obtain the content quality reflected by the content resource, that is, the target resource quality detection service is a refinement of the target service on the resource quality detection dimension. In some embodiments, the target resource quality detection service is also referred to as the downstream service of the target service.

[0119] It can be understood that compared with the first and second samples, the data volume of the third sample is very small. Therefore, through the above technical solution, only a small amount of labeled data corresponding to the target resource quality detection business is required to fine-tune the content resource detection model with good performance for downstream businesses based on the first detection model, which greatly shortens the model development time and effectively improves the efficiency of model training.

[0120] In the above technical solution, a basic model supporting multiple CV tasks in the target domain is trained based on data from the target domain. The basic model is then trained with data from the target business to be more specific in the business scenario, resulting in a first detection model. This effectively improves the generalization of the first detection model while ensuring model performance. By training the first detection model with a small amount of labeled data corresponding to the target resource quality detection task, a content resource detection model corresponding to the target resource quality detection task can be quickly obtained. The model training method is refined layer by layer, from the broad domain to the business scenario and then to the specific downstream business. The model training process fully utilizes the different samples corresponding to each level, greatly improving the accuracy of content resource detection.

[0121] The above steps 201 to 203 are a brief introduction to the training method of the content resource detection model provided in the embodiment of the present application. Next, the training method of the content resource detection model provided in the embodiment of the present application will be described in detail in combination with some embodiments. Figure 3 This is a flowchart of a method for training a content resource detection model provided in an embodiment of the present application, which is applied to the above-mentioned content resource detection system 100 and executed by the server 110. Figure 3 , methods include:

[0122] 301 . The server obtains a first sample, which includes content resources associated with keywords in a target field, and executes step 303 .

[0123] In the embodiment of the present application, the definitions of the first sample, target domain, keywords and content resources refer to step 201 and are not described in detail here.

[0124] In some embodiments, the first sample is obtained in the following ways, including but not limited to:

[0125] (1) The server searches for keywords corresponding to the target domain and obtains image samples associated with the keywords. For example, images are found through a search engine based on keywords. Since the image samples are associated with keywords, the keywords can be used to label the image samples to indicate a type of feature of the image samples. Based on this, the image samples can be used as weakly supervised samples to train the model's ability to learn image features.

[0126] In other embodiments, the keyword can also be used to construct positive image samples, that is, multiple image samples retrieved based on the same keyword can be used as a group of positive image samples.

[0127] (2) The server obtains an existing sample dataset, which includes a large amount of labeled data. For example, the ImageNet dataset or the OpenImage dataset. ImageNet is a publicly available large-scale image recognition database used by CV system recognition projects in this field, and includes a large number of labeled classified images.

[0128] It can be understood that the acquisition method of the above-mentioned multiple first samples is simple, the scale of the acquired samples is large, and there is no need to spend a lot of manpower costs to label the samples later. This not only effectively reduces the R&D cost, but also enables the first sample to bring benefits to the model performance in many aspects.

[0129] 302. The server obtains a second sample corresponding to the target business in the target domain, where the target business corresponds to the target computer vision task. The second sample includes an image sample and an image-text sample pair corresponding to the target business, where the image-text sample pair includes an image and text associated with the image, and executes step 304.

[0130] In the embodiment of the present application, the definitions of the target business, the second sample, and the target computer vision task are referred to in step 202 and will not be described in detail here.

[0131] In some embodiments, the second sample is obtained in the following ways, including but not limited to:

[0132] (1) The server determines a frame sample based on the frame pictures obtained from the content resource of the target service. The frame sample includes a positive sample constructed based on multiple similar frame pictures and a negative sample constructed based on multiple dissimilar frame pictures. It can be understood that the frame picture is a sampling result of the picture or video in the content resource. Therefore, the frame picture represents the content of the content resource to a certain extent. For example, the frame picture is a video frame extracted from the video included in the content resource, and each video frame can represent a part of the content of the video.

[0133] In some embodiments, the framed images from the same content resource can be considered as a pair or a group of similar framed images, thus providing similar image feature instances for model training. The model performs comparative learning based on positive samples, effectively training its fault-tolerant learning ability for similar images, and effectively improving the accuracy of content resource detection. Correspondingly, the framed images from different content resources can be considered as a pair or a group of dissimilar framed images, thus providing dissimilar image feature instances for model training. The model performs comparative learning based on negative samples, effectively training its ability to distinguish dissimilar images, and effectively improving the accuracy of content resource detection.

[0134] In some embodiments, the server constructs negative samples based on a temporal randomness principle and, during the construction process, deduplicates content resources, thereby removing potentially similar content resources to ensure that the frames in the negative samples are not similar. The embodiments of this application do not limit the implementation method of deduplication.

[0135] In some embodiments, the content resources of the target service may be screened based on the quality requirements of the target service to obtain the content resources corresponding to the frame-sampling video, so as to improve the quality of the frame-sampling sample pairs.

[0136] (2) Based on the images and texts included in the content resources of the target business, obtain image and text sample pairs.

[0137] In some embodiments, content resources include images and text, such as multiple images in an atlas and its description, or a video cover image and its title. In some embodiments, the image and text included in the image-text sample pair are semantically related. Therefore, the image-text sample pair can serve as a weakly supervised sample to train the model's ability to learn the semantic relationship between images and text, that is, the semantic alignment between images and text.

[0138] In some embodiments, in target scenarios corresponding to target businesses, there is no need to manually label content resources; instead, image-text sample pairs can be directly collected from the target business's content resources. For example, on a short video platform, the title text of a short video and its selected and uploaded cover image can be directly collected as a pair of associated image-text sample pairs.

[0139] It should be noted that the content resources in this application are obtained with full authorization. For example, when the producer of the short video selects the "Agree to publish" option on the short video upload page, the server obtains the specified short video that the producer agrees to publish and upload.

[0140] In some embodiments, the server preprocesses the second sample. The preprocessing includes but is not limited to: (1) image preprocessing, such as rotation, scaling, inversion, and noise addition; (2) text preprocessing, such as word segmentation, word deletion, word supplementation, and sentence splicing; (3) sequence preprocessing, such as randomly masking a portion of the image slice sequence obtained by image segmentation. By preprocessing the second sample, the characteristics of the second sample can be enhanced, the efficiency of subsequent content resource detection model training can be improved, and the accuracy of content resource detection can be improved.

[0141] It should be noted that the above step 301 can be performed before step 302 or after step 302, and this embodiment of the present application does not limit this.

[0142] Through the above technical solution, large-scale data in the target field is obtained from multiple sources as the first sample, providing a data foundation for the target field for model training; data corresponding to the target business is obtained from multiple sources to ensure the representativeness of the second sample for the target business. Therefore, at different stages of model training, the first sample and the second sample can provide effective data support, laying the foundation for improving the accuracy of content resource detection.

[0143] 303. The server trains a basic model corresponding to the target domain based on the first sample, where the basic model includes multiple types of computer vision tasks associated with the target domain.

[0144] This step refers to step 201 and is not described in detail here.

[0145] In some embodiments, the base model of the target domain is constructed based on a universal backbone model. It is understood that the universal backbone model is used to extract features of samples, and the base model can be constructed based on multiple universal backbone models to construct a representation network for extracting features of multiple types of samples. Optionally, the universal backbone model includes but is not limited to any of the following:

[0146] Deep residual network (ResNet), vision transformation model (VisionTransformer, ViT), efficient scaling model (EfficientNet), pyramid visual attention mechanism model (PyramidVision Transformer, PVT), convolutional bias visual self-attention mechanism model (Convolutional VisionTransformer, ConvViT), ResNext model sliding window hierarchical visual transformation model (Swin Transformer), data-efficient image transformation model (Data-efficient image Transformers, DeiT), bottleneck layer transformation model model (Bottleneck Transformer Network, BoT).

[0147] Through the above technical solution, multiple types of tasks can be trained in parallel. During the training process, different types of CV tasks can promote each other, realizing multi-dimensional parallel model iterative optimization, which greatly improves the generalization of the basic model and thus improves the accuracy of content resource detection.

[0148] 304. The server determines a supervision method corresponding to the second sample based on the second sample.

[0149] The corresponding relationship between the second sample and the supervision method is referred to in step 202 and will not be described in detail here.

[0150] In some embodiments, the type of the second sample determines the ability of the model to learn from it. Therefore, by adopting different supervised learning modes, the learning efficiency of the second sample can be specifically improved, thereby improving the performance of the first detection model.

[0151] 305. The server determines a target loss function corresponding to the target computer vision task based on the target business.

[0152] The definition of the target computer vision task is referred to in step 202 and will not be described in detail here.

[0153] In some embodiments, a target loss function suitable for the target computer vision task is selected based on the type of the target computer vision task. The following examples illustrate target loss functions corresponding to different types of computer vision tasks:

[0154] (1) If the target computer vision task is a classification task or a recognition task, the target loss function can be a reconstruction loss function (Reconstruction Loss) or a simulation loss function (Mimic Loss).

[0155] Among them, the reconstruction loss function is used to enable the model to learn the ability to predict the missing parts of the image.

[0156] In some embodiments, the reconstruction loss function is used to calculate the reconstruction error between the model output and the model input. Based on this reconstruction loss function, the model can predict the content of blurred areas of an image based on the image background. This allows the model to not only understand the image content but also to reconstruct reasonable hypotheses for missing portions of the image. For example, by using the surrounding context pixels to predict the image pixels in the missing portion. In some embodiments, this reconstruction loss function can implement pixel-level reconstruction loss calculation.

[0157] The mimic loss function is used to enable the model to learn the parameters of the base model. Mimic is a method for model miniaturization. This involves taking a pre-trained model, fixing its weights, and then designing a small model to learn the parameters (e.g., mimic loss) output by the pre-trained model. The pre-trained model can learn effective generalization information and pass it to the small model, allowing the small model to achieve good performance without complex data processing.

[0158] (2) The target computer vision task is a segmentation task or a detection task, and the target loss function can be a dense contrast loss function (Dense Contrastive Loss);

[0159] In some embodiments, the dense contrast loss function is used to enable the model to learn the pixel-level features of the image. It can be understood that the purpose of image classification and recognition is to determine a category for the input image, which focuses on the overall characteristics of the image; while the purpose of image target detection and semantic segmentation is to determine the part of the input image that needs attention. Therefore, the input image needs to be classified and regressed pixel by pixel. For example, in the semantic segmentation process, a category is assigned to each pixel; the purpose of target detection is to predict the category of the target part in the image and the border of the target part. Therefore, for segmentation tasks or detection tasks, it is necessary to use a dense contrast loss function to calculate the loss function at the pixel level.

[0160] This embodiment of the present application provides a schematic diagram of dense contrastive learning, see Figure 4 , where x q and x k It is a pair of positive samples obtained by transforming the same image.

[0161] In some embodiments, the server determines the target loss function based on the supervision method corresponding to the second sample and the target computer vision task.

[0162] In some embodiments, when the second sample is an image sample, the server uses a self-supervised contrastive learning method to train the basic model, and the target computer vision task is a classification task or a detection task. In this case, the target loss function can be a self-supervised contrastive loss function (Self-Contrastive Loss), which is used to enable the model to learn the features of the image. By adopting the self-supervised contrastive loss function, the model can learn how to combine the relevant features of the target business and the relevant features of other businesses from the second sample based on the self-supervised contrastive learning method, thereby obtaining a highly generalized feature representation capability.

[0163] In some embodiments, the self-supervised contrast loss function is used to compare the image and the transformed image, so that the training model can regard the image and the transformed image as the same image (or similar images), thereby enabling the model to grasp the essential features of the image.

[0164] This embodiment of the present application provides an expression for a self-supervised contrast loss function, see formula (1):

[0165]

[0166] Where L is the self-supervised contrast loss function; a n and b n are two sample features; d=||a n -b n || 2 , d represents the Euclidean distance between the two sample features of the input model; y is the label of whether the two samples match, y = 1 represents that the two samples are similar or matched, y = 0 represents that the two samples are dissimilar or mismatched; margin is the set threshold. This self-supervised contrast loss function can be used to reduce the dimensionality of sample features. The purpose is to ensure that two similar samples remain similar in the feature space after dimensionality reduction (feature extraction); and two dissimilar samples remain dissimilar in the feature space after dimensionality reduction. The embodiment of the present application provides a schematic diagram of a self-supervised contrast loss function, see Figure 5, the encoder is used to process positive samples, and the momentum encoder is used as a dynamic dictionary query to process negative samples. The momentum encoder is updated based on the parameters of the encoder.

[0167] In some embodiments, when the second sample is a picture-text sample pair, the server uses a weakly supervised contrastive learning method to train the basic model, and the target loss function is a weakly supervised contrastive loss function, which is used to enable the model to learn the similarity between the picture and the text. Among them, weakly supervised learning includes three types: (1) Incomplete Supervision: the training data includes some labeled data and some unlabeled data; (2) Inexact Supervision: the training data only has coarse-grained labels. For example, for a data packet, it is only known that the label of the data packet is Y or N, but the label of each data instance cannot be known; (3) Inaccurate Supervision: the label of the training data is not necessarily correct.

[0168] In some embodiments, the association between text and images in the image-text sample pairs may be incomplete, inaccurate or imprecise, which is consistent with the scenario targeted by weakly supervised learning. Therefore, by adopting the weakly supervised contrastive learning method, there is no need to spend money to obtain more accurate annotation data, and the model training based on the image-text sample pairs can be performed. Furthermore, based on the image-text sample pairs, the model can be trained to align images and texts in the semantic space. Optionally, a bidirectional encoder representation (Bidirectional Encoder Representation from Transformers, BERT) model can be used to extract semantic features, and pre-training learning can be performed by using contrastive loss that aligns text and images.

[0169] 306. The server trains the basic model based on the second sample, the target computer vision task, and the target loss function using a supervision method corresponding to the second sample to obtain a first detection model.

[0170] In some embodiments, the second sample is a picture, and the target loss function is a self-supervised contrast loss function. The process of training the basic model includes: inputting the picture into the basic model and performing the target computer vision task; based on the first output result of the target computer vision task, calculating the self-supervised contrast loss function and determining a first loss value, wherein the first loss value represents the error of the first detection model in learning the features of the picture; and training the first detection model based on the first loss value.

[0171] For example, a Swin Transformer model can be selected as a universal backbone model of the base model to perform image classification and recognition tasks. The base model is trained using a reconstruction loss function. The training process includes: inputting the image into the Swin Transformer model of the base model to perform the image recognition task; calculating the reconstruction loss function based on the output of the Swin Transformer model to determine the reconstruction loss value, which represents the error in the Swin Transformer model learning the features of the image; and training the Swin Transformer model based on the reconstruction loss value. It can be understood that training the Swin Transformer model is also training the first detection model.

[0172] In other embodiments, the second sample is a picture-text sample pair, and the target loss function is a weakly supervised contrast loss function. The process of training the basic model includes: inputting the picture-text sample pair into the basic model to perform the target computer vision task; based on the second output result of the target computer vision task, calculating the weakly supervised contrast loss function to obtain a second loss value, and the second loss value represents the error of the first detection model in predicting the similarity between the picture and text in the picture-text sample pair; and training the first detection model based on the second loss value.

[0173] In order to facilitate understanding of the relationship between the above-mentioned different types of second samples, different types of CV tasks and target loss functions, the present embodiment provides a schematic diagram of training the first detection model, see Figure 6 When the second sample is a picture and the CV task is a classification task or a recognition task, the self-supervised contrast learning method is adopted to train the model based on the self-supervised contrast loss function, the reconstruction loss function or the simulation loss function; when the second sample is a picture and the CV task is a detection task or a segmentation task, the self-supervised contrast learning method is adopted to train the model based on the dense contrast loss function; when the second sample is a picture-text sample pair, the weakly supervised contrast learning method is adopted to train the model based on the weakly supervised contrast loss function.

[0174] Through the above technical solution, the corresponding supervision method is selected according to the type of the second sample, and the appropriate loss function is selected based on the target CV task involved in the target business, which can effectively improve the performance of the first detection model and thus improve the accuracy of content resource detection.

[0175] In some embodiments, the image and text in the image-text sample pair can be input into different general backbone models for processing, thereby specifically extracting the image features of the image and the semantic features of the text. Then, contrastive learning under a weakly supervised contrast loss function is performed based on the image features and semantic features. Optionally, the text in the image-text sample pair can be extracted using a Bert model to extract semantic features, which is not limited in this embodiment of the present application.

[0176] For ease of understanding, the present application embodiment provides a schematic diagram of processing image and text sample pairs, see Figure 7 Among them, the text and image in the image-text sample pair are respectively extracted based on different loss functions, and then weakly supervised comparative learning is performed based on the semantic features and image features extracted respectively.

[0177] In some embodiments, there is a modality missing problem in content resources. For example, a short video only has a cover image but no title. In this example, it is necessary to simulate the situation of modality missing of content resources by constructing sample content with modality missing. For example, when processing multiple image-text sample pairs, the random probability is set to 50%, and the image-text sample pairs are randomly replaced with padding sample pairs of missing text or missing pictures. In actual training experiments, the random probability is set to 50%, which can bring better results. Through the above technical solution, during the model training process, simulating the modality missing of content resources can effectively improve the robustness of the first detection model, thereby improving the accuracy of content resource detection.

[0178] 307. The server obtains a third sample corresponding to a target resource quality detection service under the target service and a label of the third sample. The third sample includes content resources labeled based on a resource quality detection dimension corresponding to the target resource quality detection service.

[0179] In the embodiment of the present application, the definition of the target resource quality detection service, the third sample and the resource quality detection dimension can be found in Figure 1 The description of the corresponding content resource detection system is omitted here.

[0180] This step refers to step 203 and is not described in detail here.

[0181] 308. The server trains the first detection model based on the third sample and the label of the third sample to obtain a content resource detection model corresponding to the target resource quality detection service.

[0182] The principle of this step is similar to step 203 and will not be described in detail here.

[0183] In the above technical solution, a basic model supporting multiple CV tasks in the target domain is trained based on data from the target domain. The basic model is then trained with data from the target business to be more specific in the business scenario, resulting in a first detection model. This effectively improves the generalization of the first detection model while ensuring model performance. By training the first detection model with a small amount of labeled data corresponding to the target resource quality detection task, a content resource detection model corresponding to the target resource quality detection task can be quickly obtained. The model training method is refined layer by layer, from the broad domain to the business scenario and then to the specific downstream business. The model training process fully utilizes the different samples corresponding to each level, greatly improving the accuracy of content resource detection.

[0184] Furthermore, only a small number of samples corresponding to the target resource quality detection business are needed to fine-tune a content resource detection model with good performance based on the first detection model, effectively shortening the model development time and effectively improving the efficiency of model training.

[0185] In order to facilitate understanding of the above steps, the present application embodiment provides a schematic diagram of a training method for a content resource detection model, see Figure 8 , build a basic model of the target domain based on the general backbone model, and then train the basic model of the target domain based on the first sample corresponding to the target domain; train the basic model based on the second sample corresponding to the target business to obtain a first detection model; and then train the first detection model according to the third sample corresponding to the target resource quality detection business to obtain the content resource detection model of the target resource quality detection business. The principle refers to the above embodiment and will not be repeated here.

[0186] Based on the application of business scenarios, the embodiment of the present application provides a schematic diagram of a training method for a content resource detection model, see Figure 9 Among them, the basic algorithm engineer builds the basic model of the target field based on the general backbone model, and then trains the basic model of the target field based on large-scale samples. Then, through domain adaptation and model distillation, combined with the weakly supervised samples and unsupervised samples corresponding to the target business, the basic model is trained to obtain the first detection model corresponding to the target business; the business algorithm engineer, for the downstream specific business corresponding to the target resource quality detection business, based on a small number of samples of the target resource quality detection business, performs small sample transfer learning on the first detection model to obtain the content resource detection model corresponding to the target resource quality detection business.

[0187] Next, based on the above-mentioned content resource detection system and content resource detection model training method, combined with some embodiments, the content resource detection method provided in the embodiments of the present application will be introduced. Figure 10This is a flow chart of a content resource detection method provided in an embodiment of the present application, which is executed by the server in the above-mentioned content resource detection system. Figure 10 , the method comprising:

[0188] 1001. The server obtains content resources to be detected in the target domain.

[0189] The definition of the content resource to be detected refers to Figure 1 For the corresponding content and the definition of the target domain, please refer to step 201 and will not be elaborated here.

[0190] In some embodiments, the content resource to be detected is a target domain (refer to Figure 1 In some embodiments, the content resource to be detected may also be a content resource stored in the server, such as a historical short video in a short video information stream. The embodiment of the present application does not limit the method for obtaining the content resource to be detected.

[0191] 1002. The server obtains a content resource detection model corresponding to a target resource quality detection service under a target service in the target domain, and detects the content resource to be detected based on the content resource detection model. The target resource quality detection service corresponds to a resource quality detection dimension of the content resource.

[0192] The definitions of the target business, target resource quality detection business, resource quality detection dimensions, and the content resource detection model are as follows: Figure 3 The corresponding embodiments are not described in detail here.

[0193] In an embodiment of the present application, the server detects the content resources to be detected based on the content resource detection model corresponding to the target resource quality detection service to filter out low-quality content resources. For example, in a short video platform, the background content distribution server detects the acquired short videos to be distributed and filters out low-quality short videos.

[0194] In the above technical solution, based on the content resources corresponding to the target resource quality detection business, the first detection model with good overall generalization ability for the target business is fine-tuned to obtain a content resource detection model with good performance for the target resource quality detection business, thereby greatly improving the accuracy of content resource detection based on the content resource detection model.

[0195] In some embodiments, the above-mentioned content resource detection system can serve as an intermediate node connecting the content resource production end and the content resource consumption end, and is used to detect the content resources obtained from the content resource production end to filter out low-quality content resources, and then distribute the detected content resources to the content resource consumption end. This embodiment of the application provides a schematic diagram of the content resource distribution process, see Figure 11 .

[0196] In some embodiments, the content resource production end corresponds to a platform for uploading content resources to a production object of the content resource, such as a short video application. In this example, the content resource production end uploads the content resources it produces by interacting with a content resource detection system. For example, the content resource detection system includes a server. The short video creator obtains the interface address of the server through his terminal and uploads the short video he created to the server.

[0197] In some embodiments, the content resource production end may produce content resources based on any of the following content models: Professional Generated Content (PGC), User Generated Content (UGC), or Professional User Generated Content (PUGC). In other embodiments, the content resource production end may produce content resources by integrating the above-mentioned multiple models in a Multi-Channel Network (MCN) model, which is not limited in this embodiment of the present application.

[0198] In some embodiments, the content resource consumption end refers to a platform used by the consumer of the content resource to obtain the content resource, such as a short video application. In this example, the content resource consumption end interacts with the content resource detection system to obtain the detected content resource. For example, the content resource detection system includes a server, which obtains the interface address of the server based on the short video application running on the terminal, downloads the video from the server, or browses the detected short video online through the server. Optionally, the content resource consumption end can obtain the corresponding content resource from the content resource detection system based on index information. For example, the index information is a content tag for a certain type of content resource.

[0199] In some embodiments, the content resource detection system can obtain behavioral data related to content resources during the interaction with the content resource production end and the content resource consumption end, such as the reading speed and reading time of the consumer for the content resource, the loading time of the content resource, the freeze situation of the content resource, the playback click situation of the content resource, etc.

[0200] Through the above technical solution, the content resource detection system is applied in the actual content resource distribution scenario, and based on the high accuracy of content resource detection, the efficiency of content resource distribution is effectively improved.

[0201] Figure 12 This is a structural diagram of a training device for a content resource detection model provided in an embodiment of the present application, see Figure 12 , the device comprises:

[0202] A first training module 1201 is configured to train a basic model for a target domain based on a first sample corresponding to the target domain, the first sample including content resources associated with keywords in the target domain, and the basic model including multiple computer vision tasks associated with the target domain;

[0203] A second training module 1202 is configured to train the basic model based on a second sample corresponding to a target business in the target domain, using a supervision method corresponding to the second sample, to obtain a first detection model corresponding to the target business, wherein the target business corresponds to a target computer vision task, and the second sample includes an image sample and an image-text sample pair corresponding to the target business, wherein the image-text sample pair includes an image and text associated with the image;

[0204] The third training module 1203 is used to train the first detection model based on the third sample corresponding to the target resource quality detection business under the target business and the label of the third sample, to obtain a content resource detection model corresponding to the target resource quality detection business, and the third sample includes content resources annotated based on the resource quality detection dimension corresponding to the target resource quality detection business.

[0205] In one possible implementation, the second training module 1202 includes:

[0206] A first training unit is configured to, when the second sample is an image sample, train the basic model using a self-supervised contrastive learning method based on the second sample to obtain the first detection model;

[0207] The second training unit is used to train the basic model by adopting a weakly supervised contrast learning method based on the second sample to obtain the first detection model when the second sample is a picture-text sample pair.

[0208] In one possible implementation, the second training module 1202 includes:

[0209] A loss function determining unit, configured to determine a target loss function corresponding to the target computer vision task based on the target business;

[0210] The first training unit is configured to, when the second sample is an image sample, train the basic model using the self-supervised contrastive learning method based on the second sample, the target computer vision task, and the target loss function to obtain the first detection model;

[0211] The second training unit is used to train the basic model using the weakly supervised contrast learning method based on the second sample, the target computer vision task and the target loss function to obtain the first detection model when the second sample is a picture-text sample pair.

[0212] In one possible implementation, the first training unit is used to:

[0213] Input the image sample into the basic model to perform the target computer vision task;

[0214] Based on the first output result of the target computer vision task, using a self-supervised contrastive learning method, calculating the target loss function and determining a first loss value, where the first loss value represents an error in learning the features of the image sample by the first detection model;

[0215] Based on the first loss value, the first detection model is trained.

[0216] In one possible implementation, the second training unit:

[0217] Input the image-text sample pair into the basic model to perform the target computer vision task;

[0218] Based on the second output result of the target computer vision task, using a weakly supervised contrastive learning method, calculating the target loss function to obtain a second loss value, where the second loss value represents an error in the first detection model's prediction of the similarity between the image and the text in the image-text sample pair;

[0219] The first detection model is trained based on the second loss value.

[0220] In one possible implementation, if the target computer vision task is a classification task or a recognition task, the target loss function is a reconstruction loss function or a simulation loss function;

[0221] The reconstruction loss function is used to enable the first detection model to learn the ability to predict the missing part of the image; the simulation loss function is used to enable the first detection model to learn the parameters of the base model;

[0222] If the computer vision task is a segmentation task or a detection task, the objective loss function is a dense contrast loss function;

[0223] The dense contrast loss function is used to enable the first detection model to learn pixel-level features of the image.

[0224] In one possible implementation, the second sample includes:

[0225] A picture sample determined based on a frame-drawing picture obtained from a content resource of the target service, the picture sample including a positive picture sample constructed based on a plurality of similar frame-drawing pictures, and a negative picture sample constructed based on a plurality of dissimilar frame-drawing pictures;

[0226] The image-text sample pairs are determined based on the images and texts included in the content resources of the target business.

[0227] In one possible implementation, the basic model of the target domain is constructed based on a general backbone model, which includes but is not limited to any of the following:

[0228] Deep residual model, visual transformation model, pyramid visual transformation model, convolutional biased visual transformation model or sliding window layered visual transformation model.

[0229] The above technical solution uses data from the target domain as a foundation to train a basic model that supports multiple CV tasks in that domain. Data from the target business is then used to train the basic model specifically for business scenarios, resulting in a first detection model. This effectively improves the generalization of the first detection model while ensuring model performance. By training the first detection model with a small amount of labeled data corresponding to the target resource quality detection task, a content resource detection model corresponding to the target resource quality detection task can be quickly obtained. The model training method is refined layer by layer, from broad domains to business scenarios and then to specific downstream businesses. The model training process fully utilizes the different samples corresponding to each level, greatly improving the accuracy of content resource detection.

[0230] It should be noted that the training device for the content resource detection model provided in the above embodiment only uses the division of the above functional modules as an example when executing the corresponding steps. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the content resource detection model provided in the above embodiment and the training method embodiment of the content resource detection model are of the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0231] Figure 13 This is a structural diagram of a content resource detection device provided in an embodiment of the present application, see Figure 13 , the device comprises:

[0232] Acquisition module 1301, used to acquire the content resources to be detected in the target domain;

[0233] Detection module 1302 is used to obtain a content resource detection model corresponding to a target resource quality detection service under a target service under the target domain, and to detect the content resource to be detected based on the content resource detection model. The target resource quality detection service corresponds to a resource quality detection dimension of the content resource, and the content resource detection model is obtained based on the training method of the above-mentioned content resource detection model.

[0234] In the above technical solution, based on the content resources corresponding to the target resource quality detection business, the first detection model with good overall generalization ability for the target business is fine-tuned to obtain a content resource detection model with good performance for the target resource quality detection business, thereby greatly improving the accuracy of content resource detection based on the content resource detection model.

[0235] It should be noted that the content resource detection device provided in the above embodiment only uses the division of the above functional modules as an example when executing the corresponding steps. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the content resource detection device provided in the above embodiment and the content resource detection method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0236] The present embodiment provides a computer device for executing the above-mentioned content resource detection model training method or content resource detection method. The computer device herein can be implemented as the above-mentioned server 110. The structure of the computer device is described below:

[0237] Figure 14 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device 1400 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1401 and one or more memories 1402, wherein the one or more memories 1402 store at least one computer program, and the at least one computer program is loaded and executed by the one or more processors 1401 to implement the methods provided in the above-mentioned various method embodiments. Of course, the computer device 1400 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The computer device 1400 may also include other components for implementing device functions, which will not be described in detail here.

[0238] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program. The computer program can be executed by a processor to implement the content resource detection model training method or the content resource detection method in the above embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0239] In an exemplary embodiment, a computer program product or computer program is also provided, which includes a program code, which is stored in a computer-readable storage medium. The processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above-mentioned content resource detection model training method or content resource detection method.

[0240] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0241] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for training a content resource detection model, characterized in that: The method comprises: Training a basic model of the target domain based on a first sample corresponding to the target domain, wherein the first sample includes content resources associated with keywords of the target domain, and the basic model includes multiple computer vision tasks associated with the target domain; Based on a second sample corresponding to a target business in the target domain, the basic model is trained using a supervision method corresponding to the second sample to obtain a first detection model corresponding to the target business, wherein the target business corresponds to a target computer vision task, and the second sample includes an image sample and an image-text sample pair corresponding to the target business, wherein the image-text sample pair includes an image and text associated with the image; Based on the third sample corresponding to the target resource quality detection business under the target business and the label of the third sample, the first detection model is trained to obtain a content resource detection model corresponding to the target resource quality detection business, and the third sample includes content resources annotated based on the resource quality detection dimension corresponding to the target resource quality detection business.

2. The method according to claim 1, characterized in that The step of training the basic model based on a second sample corresponding to the target business in the target domain and adopting a supervision method corresponding to the second sample to obtain a first detection model corresponding to the target business includes: When the second sample is an image sample, the basic model is trained based on the second sample using a self-supervised contrastive learning method to obtain the first detection model; In the case where the second sample is a picture-text sample pair, based on the second sample, a weakly supervised contrastive learning method is adopted to train the basic model to obtain the first detection model.

3. The method according to claim 2, characterized in that The step of training the basic model based on a second sample corresponding to the target business in the target domain and adopting a supervision method corresponding to the second sample to obtain a first detection model corresponding to the target business includes: Determining a target loss function corresponding to the target computer vision task based on the target business; When the second sample is an image sample, based on the second sample, the target computer vision task, and the target loss function, the self-supervised contrastive learning method is used to train the basic model to obtain the first detection model; When the second sample is a picture-text sample pair, based on the second sample, the target computer vision task and the target loss function, the weakly supervised contrastive learning method is adopted to train the basic model to obtain the first detection model.

4. The method according to claim 3, characterized in that When the second sample is a picture sample, the self-supervised contrastive learning method is used to train the basic model based on the second sample, the target computer vision task, and the target loss function to obtain the first detection model, including: Inputting the image sample into the basic model to perform the target computer vision task; Based on the first output result of the target computer vision task, using a self-supervised contrastive learning method, calculating the target loss function and determining a first loss value, where the first loss value represents an error in learning features of the image sample by the first detection model; Based on the first loss value, the first detection model is trained.

5. The method according to claim 3, characterized in that In the case where the second sample is a picture-text sample pair, based on the second sample, the target computer vision task, and the target loss function, the weakly supervised contrastive learning method is used to train the basic model to obtain the first detection model, including: Inputting the image-text sample pair into the basic model to perform the target computer vision task; Based on the second output result of the target computer vision task, using a weakly supervised contrastive learning method, calculating the target loss function to obtain a second loss value, where the second loss value represents an error in the first detection model's prediction of the similarity between the image and text in the image-text sample pair; The first detection model is trained based on the second loss value.

6. The method according to claim 3, characterized in that If the target computer vision task is a classification task or a recognition task, the target loss function is a reconstruction loss function or a simulation loss function; The reconstruction loss function is used to enable the first detection model to learn the ability to predict missing parts of the image; the simulation loss function is used to enable the first detection model to learn the parameters of the basic model; If the computer vision task is a segmentation task or a detection task, the target loss function is a dense contrast loss function; The dense contrast loss function is used to enable the first detection model to learn pixel-level features of the image.

7. The method according to claim 1, characterized in that The second sample includes: Picture samples determined based on abstracted pictures obtained from the content resource of the target service, the picture samples including positive picture samples constructed based on a plurality of similar abstracted pictures, and negative picture samples constructed based on a plurality of dissimilar abstracted pictures; The image-text sample pairs are determined based on the images and texts included in the content resources of the target business.

8. The method according to claim 1, characterized in that The basic model of the target domain is constructed based on a general backbone model, which includes but is not limited to any of the following: Deep residual model, visual transformation model, pyramid visual transformation model, convolutional biased visual transformation model or sliding window layered visual transformation model.

9. A content resource detection method, characterized in that: The method comprises: Obtain the content resources to be tested in the target field; Obtain a content resource detection model corresponding to the target resource quality detection business under the target business under the target field, and detect the content resources to be detected based on the content resource detection model. The target resource quality detection business corresponds to the resource quality detection dimension of the content resource, and the content resource detection model is obtained based on the training method of the content resource detection model described in any one of claims 1 to 8.

10. A content resource detection system, characterized in that: The system comprises: A first database is configured to store a first sample corresponding to a target domain and a second sample corresponding to a target business within the target domain, wherein the first sample includes content resources associated with keywords in the target domain, and the second sample includes an image sample and an image-text sample pair corresponding to the target business, wherein the image-text sample pair includes an image and text associated with the image; a second database, configured to store a third sample corresponding to a target resource quality detection service under the target service and a label of the third sample, wherein the third sample includes content resources annotated based on a resource quality detection dimension corresponding to the target resource quality detection service; The server is configured to train a basic model for the target domain based on the first sample, the basic model including multiple computer vision tasks associated with the target domain; train the basic model based on a second sample corresponding to a target business under the target domain using a supervision method corresponding to the second sample to obtain a first detection model corresponding to the target business, the target business corresponding to a target computer vision task; and train the first detection model based on a third sample corresponding to a target resource quality detection business under the target business and a label of the third sample to obtain a content resource detection model corresponding to the target resource quality detection business; The server is configured to obtain the content resources to be detected in the target domain; A third database is used to store the content resource to be detected and metadata of the content resource to be detected, wherein the metadata includes relevant information of the content resource to be detected; The server is configured to schedule the content resource detection model to detect the content resource to be detected based on the third database.

11. A training device for a content resource detection model, characterized in that: The device comprises: A first training module is configured to train a basic model of a target domain based on a first sample corresponding to the target domain, wherein the first sample includes content resources associated with keywords of the target domain, and the basic model includes multiple computer vision tasks associated with the target domain; a second training module, configured to train the basic model based on a second sample corresponding to a target business in the target domain, using a supervision method corresponding to the second sample, to obtain a first detection model corresponding to the target business, wherein the target business corresponds to a target computer vision task, and the second sample includes an image sample and an image-text sample pair corresponding to the target business, wherein the image-text sample pair includes an image and text associated with the image; The third training module is used to train the first detection model based on the third sample corresponding to the target resource quality detection business under the target business and the label of the third sample, so as to obtain the content resource detection model corresponding to the target resource quality detection business, wherein the third sample includes content resources annotated based on the resource quality detection dimension corresponding to the target resource quality detection business.

12. A content resource detection device, characterized in that: The device comprises: The acquisition module is used to obtain the content resources to be detected in the target field; A detection module is used to obtain a content resource detection model corresponding to the target resource quality detection business under the target business, and detect the content resource to be detected based on the content resource detection model. The target resource quality detection business corresponds to the resource quality detection dimension of the content resource, and the content resource detection model is obtained based on the training method of the content resource detection model described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Business model training method, obstacle detection method and device, and electronic device

    CN110865421A

  • Text detection model training method and device and text detection method and device

    CN112818975A