Media asset content auditing method and device based on deep learning and medium

By annotating the historical and public data of the media system and expanding the sample set, deep learning models are trained, and the automation and efficiency of media content review are solved, and accurate identification and efficient management of sensitive content are achieved.

CN120277191APending Publication Date: 2025-07-08浪潮智能终端有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510392811.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The review process of media content in the prior art relies on manual methods, resulting in limited processing volume, high cost and impact on accuracy, and the AI model lacks sufficient sensitive content data for training.

Method used

By annotating the historical audit data and public data of the media system, an initial sample set is generated, and the sample set is expanded using image enhancement algorithm, the object detection model and feature extraction model are trained, and the detection library is built to realize automated audit of media data.

Benefits of technology

It realizes automated and efficient review of media data, improves the identification accuracy and detection efficiency of sensitive content, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277191A_ABST
    Figure CN120277191A_ABST
Patent Text Reader

Abstract

The invention discloses a medium asset content auditing method and device based on deep learning and a medium, and belongs to the technical field of data auditing. The method comprises the following steps: labeling sensitive contents in historical auditing data and public data of a media asset system to determine an initial sample set; processing a negative sample set in the initial sample set based on a preset image enhancement algorithm to generate an expanded sample set; training a preset artificial intelligence model based on the expanded sample set to obtain a media asset detection model; extracting negative feature vectors of the expanded negative sample set based on a media asset detection model, and associating the negative feature vectors with the subclass tags to generate a detection library; and inputting the to-be-audited media asset data to the media asset detection model to obtain an audit report, and feeding back the audit report to the media asset system. Through the method, the media asset data can be automatically and efficiently audited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data review, and in particular to a media content review method, device and medium based on deep learning. Background Art

[0002] The media asset management system (media asset system) is a general term for systems that manage digital storage, catalog management, retrieval and query, material transcoding, and information release of various media materials, including video, audio, text, and graphics. Its core concept is to use digital technology to manage the existing and future assets of the media, and it has a wide range of application value in the radio, television, and film and television industries. Given that media assets usually need to be disseminated to the public, their authority and authenticity require strict management of content. At each stage of content production and production, such as content uploading, quality inspection, archiving, and relocation, a strict review process must be followed to prevent the emergence of sensitive content.

[0003] Under current technical conditions, the review process of media content involves multiple links. The traditional manual review method has a limited daily processing capacity, resulting in high labor costs, and long-term high-intensity work affects the accuracy of the review. With the advancement of artificial intelligence technology, AI technology has also been applied to the field of content review. However, to realize AI review of sensitive content, sufficient relevant data must be available. Although the media system will accumulate a certain amount of sensitive data during use, that is, data screened through manual review, the above data is still insufficient for training AI models.

[0004] Therefore, how to achieve automated and efficient review of media data has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] The embodiments of the present application provide a media content review method, device and medium based on deep learning to solve the following technical problem: how to automatically and efficiently review media data.

[0006] In a first aspect, an embodiment of the present application provides a media asset content review method based on deep learning, which is applied to a media asset system. The method is characterized in that it includes: annotating sensitive content in the historical review data and public data of the media asset system to determine an initial sample set; wherein the initial sample set includes a positive sample set and a negative sample set, the positive sample set is labeled with a major category label, and the negative sample set is labeled with the position of the sensitive content, the size of the sensitive content, and the label of the sensitive content, and the label includes a major category label and a minor category label; processing the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set; wherein the augmented sample set includes an augmented negative sample set and a positive sample set; training a preset artificial intelligence model based on the augmented sample set to obtain a media asset detection model; wherein the media asset detection model includes an object detection model and a feature extraction model, the object detection model is used to identify the position and major category label of the sensitive content, and the feature extraction model is used to extract the feature vector of the sensitive content; extracting the negative feature vector of the augmented negative sample set based on the media asset detection model, and associating the negative feature vector with the minor category label to generate a detection library; inputting the media asset data to be reviewed into the media asset detection model to obtain a review report, and feeding back the review report to the media asset system.

[0007] In an implementation manner of the present application, annotating sensitive content in the historical review data and public data of the media asset system to determine an initial sample set specifically includes: obtaining public data based on a preset public channel; wherein the public data includes sensitive data and non-sensitive data; screening the data that fails the review in the media asset system as negative samples; labeling the sensitive data in the public data as negative samples, and labeling the non-sensitive data in the public data as positive samples; integrating multiple negative samples and multiple positive samples to generate an initial negative sample set.

[0008] In an implementation manner of the present application, processing the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set specifically includes: traversing the negative sample set in the initial sample set to extract the sensitive content regions in multiple negative samples; processing the sensitive content regions based on the image enhancement algorithm, and filling the sensitive content regions into the positive sample set to generate a new negative sample set; merging the new negative sample set with the initial sample set to form an augmented sample set.

[0009] In an implementation of the present application, a preset artificial intelligence model is trained based on the augmented sample set to obtain a media asset detection model, which specifically includes: dividing the augmented sample set into a training set, a validation set, and a test set; training the artificial intelligence model based on the training set to generate a trained artificial intelligence model; validating the trained artificial intelligence model based on the validation set to adjust the hyperparameters of the trained artificial intelligence model; testing the validated artificial intelligence model based on the test set to evaluate the performance of the trained artificial intelligence model; and outputting the artificial intelligence model with a performance greater than a preset performance threshold as the media asset detection model.

[0010] In an implementation of the present application, negative feature vectors of the augmented negative sample set are extracted based on the media asset detection model, and the negative feature vectors are associated with the subclass labels to generate a detection library, which specifically includes: extracting features of multiple negative samples in the augmented negative sample set based on the feature extraction model in the media asset detection model to obtain negative feature vectors; associating the negative feature vectors with corresponding subclass labels to generate retrieval items; and importing all the retrieval items into a preset vector database to generate a detection library.

[0011] In an implementation of the present application, the to-be-reviewed media asset data is input into the media asset detection model to obtain a review report, and the review report is fed back to the media asset system, which specifically includes: preprocessing the to-be-reviewed media asset data to generate corrected media asset data; scanning the corrected media asset data based on the object detection model in the media asset detection model to identify whether there is sensitive content in the corrected media asset data; if the major class label of the corrected media asset data is sensitive content, extracting the feature vectors of the sensitive area based on the feature extraction model; comparing the feature vectors with multiple retrieval items in the detection library, calculating the similarity between the feature vectors and the retrieval items, and determining the subclass label of the corrected media asset data; and generating a review report based on the subclass label.

[0012] In an implementation of the present application, comparing the feature vectors with multiple retrieval items in the detection library, calculating the similarity between the feature vectors and the retrieval items, and determining the subclass label of the corrected media asset data specifically includes: calculating the vector distance between the feature vectors and the negative feature vectors among the multiple retrieval items based on a preset vector distance algorithm; selecting the retrieval item corresponding to the negative feature vector with the minimum vector distance as the matching result; and determining the subclass label of the corrected media asset data based on the subclass label in the matching result.

[0013] In an implementation of the present application, the method further includes: parsing the review report to determine sensitive content; and processing the sensitive content based on preset processing rules.

[0014] In a second aspect, an embodiment of the present application further provides a media asset content review device based on deep learning. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: label sensitive content in the historical review data and public data of the media asset system to determine an initial sample set; wherein the initial sample set includes a positive sample set and a negative sample set, the positive sample set is labeled with a major category label, and the negative sample set is labeled with the location of the sensitive content, the size of the sensitive content, and the label of the sensitive content, and the label includes a major category label and a minor category label; process the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set; wherein the augmented sample set includes an augmented negative sample set and a positive sample set; train a preset artificial intelligence model based on the augmented sample set to obtain a media asset detection model; wherein the media asset detection model includes an object detection model and a feature extraction model, the object detection model is used to identify the location and major category label of the sensitive content, and the feature extraction model is used to extract the feature vector of the sensitive content; extract the negative feature vector of the augmented negative sample set based on the media asset detection model, and associate the negative feature vector with the minor category label to generate a detection library; input the media asset data to be reviewed into the media asset detection model to obtain a review report, and feedback the review report to the media asset system.

[0015] In a third aspect, an embodiment of the present application further provides a non-volatile computer storage medium for media content review based on deep learning, storing computer-executable instructions, and the computer-executable instructions are configured to: label sensitive content in the historical review data and public data of the media system to determine an initial sample set; wherein, the initial sample set includes a positive sample set and a negative sample set, the positive sample set is labeled with a major category label, and the negative sample set is labeled with the position of the sensitive content, the size of the sensitive content, and the label of the sensitive content, and the label includes a major category label and a minor category label; process the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set; wherein, the augmented sample set includes an augmented negative sample set and a positive sample set; train a preset artificial intelligence model based on the augmented sample set to obtain a media detection model; wherein, the media detection model includes an object detection model and a feature extraction model, the object detection model is used to identify the position and major category label of the sensitive content, and the feature extraction model is used to extract the feature vector of the sensitive content; extract the negative feature vector of the augmented negative sample set based on the media detection model, and associate the negative feature vector with the minor category label to generate a detection library; input the media data to be reviewed into the media detection model to obtain a review report, and feedback the review report to the media system.

[0016] A media content review method, device and medium based on deep learning provided by an embodiment of the present application generate an initial sample set including positive and negative samples by labeling sensitive content in historical review data and public data, providing a basis for subsequent model training. The sample set is augmented through an image enhancement algorithm, effectively solving the problem of insufficient sensitive content data volume and enhancing the generalization ability of the model. By training an artificial intelligence model including an object detection model and a feature extraction model, accurate identification of the position, size and label of sensitive content is achieved. Further, negative sample feature vectors are extracted based on the trained model and associated with minor category labels to construct a detection library, providing support for fast retrieval and matching. During the review process, the media data to be reviewed can be automatically scanned, sensitive content can be identified and a review report can be generated, realizing the automation of the review process. In summary, the present application can automate and efficiently review media data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a flowchart of a media content review method based on deep learning provided by an embodiment of the present application; Figure 2Schematic diagram of the internal structure of a media asset content review device provided by an embodiment of the present application. Detailed implementation manners

[0018] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0019] The embodiments of the present application provide a media asset content review method, device, and medium based on deep learning to solve the following technical problems: how to automatically and efficiently review media asset data.

[0020] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the drawings.

[0021] Figure 1 A flowchart of a media asset content review based on deep learning provided by an embodiment of the present application. As Figure 1 shown, a media asset content review method based on deep learning provided by an embodiment of the present application specifically includes the following steps: Step 1: Label sensitive content in the historical review data and public data of the media asset system to determine an initial sample set; wherein, the initial sample set includes a positive sample set and a negative sample set, the positive sample set is labeled with a major category label, and the negative sample set is labeled with the location of the sensitive content, the size of the sensitive content, and the label of the sensitive content, and the label includes a major category label and a minor category label.

[0022] By obtaining public data, screening data that fails the review as negative samples, labeling sensitive and non-sensitive data in the public data, and integrating and generating an initial sample set, the diversity of the initial sample set can be ensured to a certain extent to support subsequent model training.

[0023] Step 11: Obtain public data based on a preset public channel; wherein, the public data includes sensitive data and non-sensitive data.

[0024] Obtain public data through a preset public channel (such as short video platforms, social media, and news websites). The public data needs to include sensitive data (such as videos including disgraced artists) and non-sensitive data (such as normal entertainment videos). For example, obtain videos including disgraced artists as sensitive data from a short video platform, and the above videos include the facial features of artist Zhang; at the same time, obtain normal entertainment videos as non-sensitive data, such as dance teaching videos and food production videos.

[0025] Step 12: Screen the data that fails the review in the media asset system as negative samples.

[0026] The data that fails the review includes video clips determined by the media asset system or manual review to contain sensitive content (such as artists with bad records).

[0027] Step 13: Label the sensitive data in the public data as negative samples and the non-sensitive data in the public data as positive samples.

[0028] Label the sensitive data in the public data (including videos of artists with bad records), clarify the location (such as the timestamp in the video), size (such as the proportion of the sensitive content in the video), and label (including the major label "sensitive content" and the minor label "artist with bad record - Zhang - facial features") of its sensitive content, and use it as a negative sample. At the same time, label the non-sensitive data in the public data (such as normal entertainment videos), clarify its major label (such as "non-sensitive content - dance teaching"), and use it as a positive sample.

[0029] Step 14: Integrate multiple negative samples and multiple positive samples to generate an initial sample set.

[0030] Integrate the screened data that fails the review (negative samples) and the labeled public data (negative samples and positive samples of non-sensitive data) to generate an initial sample set. The initial sample set includes a positive sample set and a negative sample set. The positive sample set is labeled with major labels, and the negative sample set is labeled with the location, size, and label of the sensitive content.

[0031] For example, the initial sample set includes 10,000 positive samples and 1,000 negative samples. Among them, the positive samples are labeled as "non-sensitive content - dance teaching", "non-sensitive content - food making", etc., and the negative samples are labeled as "sensitive content - artist with bad record - Zhang - facial features", "sensitive content - violent scene", etc.

[0032] Step 2: Process the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set; among them, the augmented sample set includes an augmented negative sample set and a positive sample set.

[0033] Process the initial sample set through an image enhancement algorithm to generate an augmented sample set to increase the diversity and quantity of the sample set and improve the generalization ability and robustness of the model.

[0034] Step 21: Traverse the negative sample set in the initial sample set to extract the sensitive content areas in multiple negative samples.

[0035] Traverse the negative sample set in the initial sample set and extract the sensitive content areas in each negative sample. For example, in the scenario of detecting disgraced artists in short videos, traverse the video negative samples including disgraced artists and extract the areas where the disgraced artists appear in the videos, such as facial features and body movements.

[0036] Step 22: Process the sensitive content areas based on the image enhancement algorithm and fill the sensitive content areas into the positive sample set to generate a new negative sample set.

[0037] Process the extracted sensitive content areas through a preset image enhancement algorithm (such as rotation, scaling, brightness adjustment) to generate diverse sensitive content areas. Fill the processed sensitive content areas into the positive sample set (such as normal entertainment videos) to generate a new negative sample set. For example, after rotating and adjusting the brightness of the facial features of a disgraced artist, fill them into a dance teaching video to generate a new video negative sample including the disgraced artist.

[0038] Step 23: Merge the new negative sample set with the initial sample set to form an expanded sample set.

[0039] Merge the generated new negative sample set with the initial sample set to form an expanded sample set. The expanded sample set includes more negative and positive samples, improving the diversity and quantity of the sample set, which helps the model better learn the features of sensitive content and improve the detection accuracy.

[0040] In a specific example, in the scenario of detecting disgraced artists on a short video platform, the initial sample set includes 10,000 positive samples (such as dance teaching videos, food production videos) and 1,000 negative samples (such as videos including disgraced artists). By traversing the negative sample set, extract the facial feature areas of the disgraced artists in each negative sample. Then, process the above facial feature areas through an image enhancement algorithm (such as rotation, scaling, brightness adjustment) to generate diverse facial feature areas. Fill the processed facial feature areas into the positive sample set to generate new video negative samples including disgraced artists. Finally, merge the newly generated negative sample set with the initial sample set to form an expanded sample set, which includes 10,000 positive samples and 10,000 negative samples.

[0041] Step 3: Train a preset artificial intelligence model based on the expanded sample set to obtain a media asset detection model; wherein, the media asset detection model includes an object detection model and a feature extraction model, and the object detection model is used to identify the location and major category labels of sensitive content, and the feature extraction model is used to extract the feature vectors of sensitive content.

[0042] Obtain a media asset detection model through training an artificial intelligence model, which is used to identify the location, major category labels of sensitive content and extract the feature vectors of sensitive content.

[0043] Step 31: Divide the augmented sample set into a training set, a validation set, and a test set.

[0044] Divide the augmented sample set into a training set, a validation set, and a test set according to a preset ratio (such as 7:2:1). The training set is used to train the artificial intelligence model, the validation set is used to adjust the hyperparameters of the model, and the test set is used to evaluate the performance of the model. For example, in the scenario of detecting disgraced artists in short videos, the augmented sample set contains 20,000 positive samples and 20,000 negative samples, and is divided into 14,000 training samples, 4,000 validation samples, and 2,000 test samples according to the ratio of 7:2:1.

[0045] Step 32: Train the artificial intelligence model based on the training set to generate a trained artificial intelligence model.

[0046] Train a preset artificial intelligence model through the training set to generate a trained artificial intelligence model. During the training process, the model learns to identify the positions of sensitive content, the major category labels, and extract the feature vectors of sensitive content. For example, in the scenario of detecting disgraced artists in short videos, the trained artificial intelligence model can identify the positions where disgraced artists appear in the video, the major category labels (such as people), and extract the feature vectors of disgraced artists.

[0047] Step 33: Validate the trained artificial intelligence model based on the validation set to adjust the hyperparameters of the trained artificial intelligence model.

[0048] Validate the trained artificial intelligence model through the validation set, evaluate the performance of the model, and adjust the hyperparameters (learning rate, number of iterations) of the model according to the validation results, so as to improve the generalization ability and robustness of the model. For example, in the scenario of detecting disgraced artists in short videos, evaluate the recognition accuracy of the model through the validation set, and adjust the learning rate and number of iterations of the model to improve the recognition accuracy of the model.

[0049] Step 34: Test the validated artificial intelligence model based on the test set to evaluate the performance of the trained artificial intelligence model.

[0050] Test the validated artificial intelligence model through the test set to evaluate the performance of the model. The test set is independent of the training set and the validation set, and objectively reflects the generalization ability and robustness of the model. Through testing, evaluate the accuracy and recall rate indicators of the model in identifying sensitive content.

[0051] Step 35: Output the artificial intelligence model with performance greater than the preset performance threshold as the media asset detection model.

[0052] Output the artificial intelligence model with performance greater than the preset performance threshold on the test set as the media asset detection model. The preset performance threshold is set according to actual needs. For example, it can be set that the accuracy rate is greater than 95% or the recall rate is greater than 90%.

[0053] In a specific example, in the scenario of detecting delinquent artists on a short video platform, the augmented sample set includes 20,000 samples, among which 14,000 are used as the training set, 7,000 are used as the validation set, and 1,000 are used as the test set. The artificial intelligence model is trained with the training set. During the training process, the object detection model learns to identify the positions and major class labels of delinquent artists, and the feature extraction model learns to extract the feature vectors of delinquent artists. Then, the trained model is verified with the validation set, and the hyperparameters of the model, such as the learning rate and the number of iterations, are adjusted according to the verification results. Finally, the verified model is tested with the test set to evaluate the performance of the model. If the accuracy rate of the model on the test set is greater than 95%, then the model is output as the media asset detection model for detecting delinquent artists on the short video platform.

[0054] Step 4: Based on the media asset detection model, extract the negative feature vectors of the augmented negative sample set, and associate the negative feature vectors with the minor class labels to generate a detection library.

[0055] Extract the negative feature vectors of the augmented negative sample set through the media asset detection model, and associate the negative feature vectors with the minor class labels to generate a detection library. The detection library can be used for quickly retrieving and matching sensitive content to improve the detection efficiency.

[0056] Step 41: Based on the feature extraction model in the media asset detection model, extract features from multiple negative samples in the augmented negative sample set to obtain negative feature vectors.

[0057] Extract features from each negative sample in the augmented negative sample set through the feature extraction model in the media asset detection model to obtain negative feature vectors. The negative feature vectors can characterize the sensitive content features in the negative samples. For example, in the scenario of detecting delinquent artists on a short video platform, the negative feature vectors characterize the facial features and body movements of delinquent artists.

[0058] Step 42: Associate the negative feature vectors with the corresponding minor class labels to generate retrieval items.

[0059] Minor class labels are used to further subdivide the types of sensitive content. For example, in the scenario of detecting delinquent artists on a short video platform, the minor class labels include "Delinquent Artist - Zhang - Facial Features" and "Delinquent Artist - Li - Facial Features". By associating the negative feature vectors with the minor class labels, the refined management of sensitive content can be achieved.

[0060] Step 43: Import all the retrieval items into a preset vector database to generate a detection library.

[0061] The vector database supports efficient vector retrieval and matching, quickly retrieving sensitive content similar to the content to be detected. The retrieval items in the detection library include negative feature vectors and subclass labels, and the matching sensitive content can be quickly found in the vector database through the retrieval items, improving the detection efficiency.

[0062] Step 5: Input the media asset data to be audited into the media asset detection model to obtain an audit report, and feedback the audit report to the media asset system.

[0063] Input the media asset data to be audited into the media asset detection model, generate an audit report through model processing, and feedback the report to the media asset system. The audit report includes the identification results of sensitive content and subclass label information in the media asset data, helping the media asset system manage and process sensitive content.

[0064] Step 51: Preprocess the media asset data to be audited to generate corrected media asset data.

[0065] The preprocessing includes data cleaning, format conversion, and resolution adjustment operations, aiming to improve the accuracy and efficiency of subsequent sensitive content identification. For example, in the scenario of detecting disgraced artists on a short video platform, the preprocessing includes converting the video to a unified resolution and frame rate for better model processing.

[0066] Step 52: Scan the corrected media asset data based on the object detection model in the media asset detection model to identify whether there is sensitive content in the corrected media asset data.

[0067] The object detection model can locate the position of sensitive content and give a major class label. For example, in a short video, identify the position of a disgraced artist and label it as "disgraced artist".

[0068] Step 53: If the major class label of the corrected media asset data is sensitive content, extract the feature vector of the sensitive area based on the feature extraction model.

[0069] The feature vector can characterize the features of sensitive content. For example, the facial features and body movements of a disgraced artist.

[0070] Step 54: Compare the feature vector with multiple retrieval items in the detection library, calculate the similarity between the feature vector and the retrieval item, and determine the subclass label of the corrected media asset data.

[0071] Compare the extracted feature vectors with multiple retrieval items in the detection library, calculate the similarity between the feature vectors and the retrieval items, and determine the subclass label of the corrected media asset data. The subclass label is used to further subdivide the types of sensitive content. For example, in the scenario of detecting delinquent artists on a short video platform, the subclass labels include "Delinquent Artist - Zhang - Facial Features" and "Delinquent Artist - Li - Facial Features". By comparing the similarity between the feature vectors and the retrieval items, refined management of sensitive content is achieved.

[0072] Step 541: Calculate the vector distance between the feature vector and the negative feature vectors among the multiple retrieval items based on a preset vector distance algorithm.

[0073] For example, the vector distance algorithm uses the cosine distance algorithm to measure the similarity between the feature vector and the negative feature vector. The smaller the vector distance, the more similar the feature vector and the negative feature vector are.

[0074] Step 542: Select the retrieval item corresponding to the negative feature vector with the smallest vector distance as the matching result.

[0075] Select the retrieval item corresponding to the negative feature vector with the smallest vector distance as the matching result. The subclass label in the matching result is the subclass label of the corrected media asset data.

[0076] Step 543: Determine the subclass label of the corrected media asset data based on the subclass label in the matching result.

[0077] The subclass label is used to further describe the characteristics of sensitive content. For example, in the scenario of detecting delinquent artists on a short video platform, the subclass labels include "Delinquent Artist - Zhang - Facial Features" and "Delinquent Artist - Li - Facial Features".

[0078] Step 55: Generate an audit report based on the subclass label.

[0079] The audit report includes the recognition result of sensitive content and subclass label information, which helps the media asset system manage and process sensitive content. For example, in the scenario of detecting delinquent artists on a short video platform, the audit report includes the location of the delinquent artist identified in the short video and subclass label information.

[0080] In the scenario of detecting delinquent artists on a short-video platform, the media asset data to be audited is a short video. First, preprocess the short video to convert it to a unified resolution and frame rate. Then, scan the short video through the object detection model in the media asset detection model to identify whether there are delinquent artists in it. If a delinquent artist is identified, extract the feature vector of the area of the delinquent artist through the feature extraction model. Next, compare the feature vector with multiple retrieval items in the detection library and calculate the similarity between the feature vector and the retrieval items. Select the retrieval item corresponding to the negative feature vector with the smallest vector distance as the matching result, and determine the subclass label of the delinquent artist in the short video based on the subclass label in the matching result. Finally, generate an audit report based on the subclass label and feedback the report to the short-video platform.

[0081] After generating the audit report, it is also necessary to process the media asset data corresponding to the audit report. Therefore, this application also includes the following method: Parse the audit report to determine sensitive content, and process the sensitive content based on preset processing rules.

[0082] By parsing the audit report, determine the sensitive content in the media asset data. For example, in the scenario of detecting delinquent artists on a short-video platform, the audit report includes the position and subclass label information of the delinquent artists identified in the short video. Then, process the sensitive content based on preset processing rules. The preset processing rules are defined according to actual needs. For example, take down, delete, or restrict the dissemination of short videos including delinquent artists.

[0083] The above is the method embodiment proposed in this application. Based on the same inventive concept, the embodiments of this application also provide a media asset content audit device based on deep learning, and its structure is as Figure 2 shown.

[0084] Figure 2 This is a schematic internal structure diagram of a media asset content audit device based on deep learning provided by the embodiment of this application. As Figure 2 shown, the device includes: At least one processor 201; And a memory 202 communicatively connected to at least one processor; Wherein, the memory 202 stores instructions executable by at least one processor, and the instructions are executed by at least one processor 201 so that at least one processor 201 can: Annotate sensitive content in the historical review data and public data of the media asset system to determine an initial sample set; wherein, the initial sample set includes a positive sample set and a negative sample set, the positive sample set is labeled with a major category label, and the negative sample set is labeled with the location of the sensitive content, the size of the sensitive content, and the label of the sensitive content, and the label includes a major category label and a minor category label; Process the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set; wherein, the augmented sample set includes an augmented negative sample set and a positive sample set; Train a preset artificial intelligence model based on the augmented sample set to obtain a media asset detection model; wherein, the media asset detection model includes an object detection model and a feature extraction model, the object detection model is used to identify the location and major category label of the sensitive content, and the feature extraction model is used to extract the feature vector of the sensitive content; Extract the negative feature vectors of the augmented negative sample set based on the media asset detection model, and associate the negative feature vectors with the minor category labels to generate a detection library; Input the media asset data to be reviewed into the media asset detection model to obtain a review report, and feedback the review report to the media asset system.

[0085] Some embodiments of the present application provide a non-volatile computer storage medium corresponding to Figure 1 for media asset content review based on deep learning, storing computer-executable instructions, and the computer-executable instructions are set as: Annotate sensitive content in the historical review data and public data of the media asset system to determine an initial sample set; wherein, the initial sample set includes a positive sample set and a negative sample set, the positive sample set is labeled with a major category label, and the negative sample set is labeled with the location of the sensitive content, the size of the sensitive content, and the label of the sensitive content, and the label includes a major category label and a minor category label; Process the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set; wherein, the augmented sample set includes an augmented negative sample set and a positive sample set; Train a preset artificial intelligence model based on the augmented sample set to obtain a media asset detection model; wherein, the media asset detection model includes an object detection model and a feature extraction model, the object detection model is used to identify the location and major category label of the sensitive content, and the feature extraction model is used to extract the feature vector of the sensitive content; Extract the negative feature vectors of the augmented negative sample set based on the media asset detection model, and associate the negative feature vectors with the minor category labels to generate a detection library; Input the media asset data to be reviewed into the media asset detection model to obtain a review report, and feedback the review report to the media asset system.

[0086] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the Internet of Things devices and media, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.

[0087] The systems and media provided by the embodiments of the present application correspond one-to-one with the methods. Therefore, the systems and media also have beneficial technical effects similar to those of the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be elaborated here.

[0088] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage) that include computer-usable program code.

[0089] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0090] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0091] The above computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.

[0092] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0093] The memory includes non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0094] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented by any method or technology for information storage. The information is computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0095] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of another identical element in the process, method, commodity or device comprising the element.

[0096] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, or improvement made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A media asset content review method based on deep learning, applied to a media asset system, characterized in that The method includes: Labeling sensitive content in the historical review data and public data of the media asset system to determine an initial sample set; wherein, the initial sample set includes a positive sample set and a negative sample set, the positive sample set is labeled with a major category label, and the negative sample set is labeled with the location of the sensitive content, the size of the sensitive content, and the label of the sensitive content, and the label includes a major category label and a minor category label; Processing the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set; wherein, the augmented sample set includes an augmented negative sample set and a positive sample set; Training a preset artificial intelligence model based on the augmented sample set to obtain a media asset detection model; wherein, the media asset detection model includes an object detection model and a feature extraction model, the object detection model is used to identify the location and major category label of the sensitive content, and the feature extraction model is used to extract the feature vector of the sensitive content; Extracting the negative feature vectors of the augmented negative sample set based on the media asset detection model, and associating the negative feature vectors with the minor category labels to generate a detection library; Inputting the media asset data to be reviewed into the media asset detection model to obtain a review report, and feeding back the review report to the media asset system.

2. The media asset content review method based on deep learning according to claim 1, wherein Labeling sensitive content in the historical review data and public data of the media asset system to determine an initial sample set, specifically including: Obtaining public data based on a preset public channel; wherein, the public data includes sensitive data and non-sensitive data; Selecting the data that fails the review in the media asset system as negative samples; Labeling the sensitive data in the public data as negative samples, and labeling the non-sensitive data in the public data as positive samples; Integrating multiple negative samples and multiple positive samples to generate an initial negative sample set.

3. The media asset content review method based on deep learning according to claim 1, characterized in that, Processing the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set, specifically including: Traversing the negative sample set in the initial sample set to extract the sensitive content areas in multiple negative samples; Processing the sensitive content areas based on the image enhancement algorithm, and filling the sensitive content areas into the positive sample set to generate a new negative sample set; Merging the new negative sample set with the initial sample set to form an augmented sample set.

4. A media asset content review method based on deep learning according to claim 1, characterized in that, Training a preset artificial intelligence model based on the augmented sample set to obtain a media asset detection model, specifically including: Dividing the augmented sample set into a training set, a validation set, and a test set; Training the artificial intelligence model based on the training set to generate a trained artificial intelligence model; Validating the trained artificial intelligence model based on the validation set to adjust the hyperparameters of the trained artificial intelligence model; Testing the validated artificial intelligence model based on the test set to evaluate the performance of the trained artificial intelligence model; Outputting the artificial intelligence model with a performance greater than a preset performance threshold as the media asset detection model.

5. A media asset content review method based on deep learning according to claim 1, characterized in that, Extracting the negative feature vectors of the augmented negative sample set based on the media asset detection model, and associating the negative feature vectors with the minor category labels to generate a detection library, specifically including: Based on the feature extraction model in the media asset detection model, extract features from multiple negative samples in the augmented negative sample set to obtain negative feature vectors; Associate the negative feature vectors with corresponding subclass labels to generate retrieval items; Import all the retrieval items into a preset vector database to generate a detection library.

6. The media asset content review method based on deep learning according to claim 1, wherein, Input the media asset data to be audited into the media asset detection model to obtain an audit report and feedback the audit report to the media asset system, which specifically includes: Preprocess the media asset data to be audited to generate corrected media asset data; Scan the corrected media asset data based on the object detection model in the media asset detection model to identify whether there is sensitive content in the corrected media asset data; If the major class label of the corrected media asset data is sensitive content, extract the feature vectors of the sensitive area based on the feature extraction model; Compare the feature vectors with multiple retrieval items in the detection library, calculate the similarity between the feature vectors and the retrieval items to determine the subclass label of the corrected media asset data; Generate an audit report based on the subclass label.

7. The media asset content review method based on deep learning according to claim 6, characterized in that, Compare the feature vectors with multiple retrieval items in the detection library, calculate the similarity between the feature vectors and the retrieval items to determine the subclass label of the corrected media asset data, which specifically includes: Calculate the vector distance between the feature vectors and the negative feature vectors among multiple retrieval items based on a preset vector distance algorithm; Select the retrieval item corresponding to the negative feature vector with the smallest vector distance as the matching result; Determine the subclass label of the corrected media asset data based on the subclass label in the matching result.

8. A media asset content review method based on deep learning according to claim 1, characterized in that, The method further includes: Parse the audit report to determine sensitive content; Process the sensitive content based on preset processing rules.

9. A media asset content review device based on deep learning, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Label sensitive content in the historical audit data and public data of the media asset system to determine an initial sample set; wherein, the initial sample set includes a positive sample set and a negative sample set, the positive sample set is labeled with a major class label, and the negative sample set is labeled with the position of sensitive content, the size of sensitive content, and the label of sensitive content, and the label includes a major class label and a subclass label; Process the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set; wherein, the augmented sample set includes an augmented negative sample set and a positive sample set; Train a preset artificial intelligence model based on the augmented sample set to obtain a media asset detection model; wherein, the media asset detection model includes an object detection model and a feature extraction model, the object detection model is used to identify the position and major class label of sensitive content, and the feature extraction model is used to extract the feature vectors of sensitive content; Extract the negative feature vectors of the augmented negative sample set based on the media asset detection model, and associate the negative feature vectors with the subclass labels to generate a detection library; Input the media asset data to be audited into the media asset detection model to obtain an audit report, and feedback the audit report to the media asset system.

10. A non-volatile computer storage medium for media content review based on deep learning, storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: Annotate sensitive content in the historical audit data and public data of the media asset system to determine an initial sample set; wherein, the initial sample set includes a positive sample set and a negative sample set, the positive sample set is annotated with a major category label, and the negative sample set is annotated with the location of the sensitive content, the size of the sensitive content, and the label of the sensitive content, and the label includes a major category label and a minor category label; Process the initial sample set based on a preset image enhancement algorithm to generate an augmented sample set; wherein, the augmented sample set includes an augmented negative sample set and a positive sample set; Train a preset artificial intelligence model based on the augmented sample set to obtain a media asset detection model; wherein, the media asset detection model includes an object detection model and a feature extraction model, the object detection model is used to identify the location and major category label of the sensitive content, and the feature extraction model is used to extract the feature vector of the sensitive content; Extract the negative feature vectors of the augmented negative sample set based on the media asset detection model, and associate the negative feature vectors with the minor category labels to generate a detection library; Input the media asset data to be audited into the media asset detection model to obtain an audit report, and feedback the audit report to the media asset system.