Content detection method, device and computer-readable storage medium
Through multimodal feature extraction and modal weighting methods, the problem of difficulty in detecting noise-infringing content in the prior art is solved, and the accuracy and efficiency of content detection are improved.
Patent Information
- Application Number
- CN202210060921.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-01-19
AI Technical Summary
The prior art is difficult to effectively detect infringing content after noise addition, resulting in poor accuracy and low efficiency of content detection results.
By obtaining the content to be detected and the source content set, multimodal feature extraction is performed, modal similarity is calculated, and weighted based on the predicted modal weights, and finally the copyright information of the content to be detected is detected in the source content set.
It improves the accuracy and efficiency of content detection results, enhances the influence of important modes in the detection content, and reduces the impact of artificially introduced noise on the detection results.
Smart Images

Figure CN114461987B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, and in particular, to a content detection method, apparatus, and computer-readable storage medium. Background Art
[0002] In recent years, with the rapid development of Internet technologies, more and more works of content such as videos, pictures, and texts have been widely published on various platforms. Among them, there is a lot of content that infringes on the legitimate rights and interests of the rights holders in the published content. These infringing contents avoid content detection through artificially introduced noises, endangering the legitimate rights and interests of the rights holders.
[0003] In the process of researching and practicing the prior art, the inventors of the present invention found that the prior art cannot well detect whether the infringing content after adding noise is infringing, resulting in poor accuracy of content detection results and low content detection efficiency. Summary of the Invention
[0004] Embodiments of this application provide a content detection method, apparatus, and computer-readable storage medium, which can improve the accuracy of content detection results and thus improve content detection efficiency.
[0005] Embodiments of this application provide a content detection method, including:
[0006] Obtain the content to be detected and a source content set, where the source content set includes at least one source content with copyright;
[0007] Perform multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and perform multi-modal feature extraction on the source content to obtain the source modal features of each modality;
[0008] Calculate the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality;
[0009] Obtain at least one content sample pair, where the content sample pair includes a detected content sample, a source content sample, and labeled copyright information;
[0010] Use the weight determination sub-model in the preset content detection model to perform multi-modal feature extraction on the detected content sample to obtain the sample modal features of each modality;
[0011] Calculate the correlation coefficient between the sample modal features of each modality and the labeled copyright information;
[0012] Based on the correlation coefficient, determine the predicted modal weight of each modality;
[0013] Predict the modal similarity of the content sample pair based on the predicted modal weights to obtain the predicted modal similarity of each modality;
[0014] Converge the preset content detection model according to the predicted modal similarity and the labeled copyright information to obtain a trained content detection model;
[0015] According to the modal features to be detected, use the trained content detection model to determine the modal weights of each modality, and weight the modal similarity based on the modal weights;
[0016] Based on the weighted modal similarity, detect the copyright information of the content to be detected in the source content set.
[0017] Correspondingly, an embodiment of the present application provides a content detection device, including:
[0018] An acquisition unit, configured to acquire the content to be detected and a source content set, where the source content set includes at least one source content with copyright;
[0019] A feature extraction unit, configured to perform multi-modal feature extraction on the content to be detected to obtain the modal features to be detected of each modality, and perform multi-modal feature extraction on the source content to obtain the source modal features of each modality;
[0020] A calculation unit, configured to calculate the similarity between the modal features to be detected and the source modal features of the corresponding modality to obtain the modal similarity of each modality;
[0021] A determination unit, configured to obtain at least one content sample pair, where the content sample pair includes a detected content sample, a source content sample, and labeled copyright information; perform multi-modal feature extraction on the detected content sample by using a weight determination sub-model in the preset content detection model to obtain the sample modal features of each modality; calculate the correlation coefficient between the sample modal features of each modality and the labeled copyright information; based on the correlation coefficient, determine the predicted modal weights of each modality; based on the predicted modal weights, predict the modal similarity of the content sample pair to obtain the predicted modal similarity of each modality; converge the preset content detection model according to the predicted modal similarity and the labeled copyright information to obtain a trained content detection model; according to the modal features to be detected, use the trained content detection model to determine the modal weights of each modality, and weight the modal similarity based on the modal weights;
[0022] A detection unit, configured to detect the copyright information of the content to be detected in the source content set based on the weighted modal similarity.
[0023] In one embodiment, the determination unit includes:
[0024] A modal importance score identification subunit for identifying the modal importance score corresponding to the to-be-detected modal feature;
[0025] A modal quality score detection subunit for detecting the modal quality score of the to-be-detected modal feature;
[0026] A coefficient fusion subunit for fusing the modal importance score and the modal quality score of the corresponding modality to obtain the modal weight of each modality.
[0027] In one embodiment, the detection unit includes:
[0028] A similarity fusion subunit for fusing the weighted modal similarities corresponding to each modality to obtain the total weighted modal similarity corresponding to each source content;
[0029] A screening subunit for screening out the target source content from the source content set according to the total weighted modal similarity;
[0030] A copyright information determination subunit for determining the copyright information of the to-be-detected content based on the target source content.
[0031] In one embodiment, the feature extraction unit includes:
[0032] A modal extraction subunit for performing multi-modal extraction on the to-be-detected content to obtain the to-be-detected modal data of each modality, and performing multi-modal extraction on the source content to obtain the source modal data of each modality;
[0033] An encoder determination subunit for respectively determining the encoders corresponding to the to-be-detected modal data and the source modal data according to the modal type corresponding to each modality;
[0034] A to-be-detected extraction subunit for performing feature extraction on the to-be-detected modal data based on the encoder corresponding to the to-be-detected modal data to obtain the to-be-detected modal feature of each modality;
[0035] A source extraction subunit for performing feature extraction on the source modal data based on the encoder corresponding to the source modal data to obtain the source modal feature of each modality.
[0036] In one embodiment, the convergence unit includes:
[0037] A calculation subunit for calculating the total predicted modal similarity of the content sample pair according to the predicted modal similarity of each modality;
[0038] A loss information determination subunit for determining the target loss information of each content sample pair based on the total predicted modal similarity and the corresponding labeled copyright information.
[0039] A convergence subunit, configured to converge the preset content detection model based on the target loss information to obtain a trained content detection model.
[0040] In addition, an embodiment of the present application further provides a computer-readable storage medium storing multiple instructions, which are suitable for being loaded by a processor to execute the steps in any content detection method provided by the embodiments of the present application.
[0041] In addition, an embodiment of the present application further provides a computer device, including a processor and a memory, where the memory stores an application program, and the processor is configured to run the application program in the memory to implement the content detection method provided by the embodiments of the present application.
[0042] An embodiment of the present application further provides a computer program product or a computer program, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the content detection method provided by the embodiments of the present application.
[0043] In the embodiment of the present application, by obtaining the content to be detected and a source content set, where the source content set includes at least one source content with copyright; performing multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and performing multi-modal feature extraction on the source content to obtain the source modal features of each modality; calculating the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality; determining the modal weight of each modality according to the to-be-detected modal features, and weighting the modal similarity based on the modal weight; detecting the copyright information of the content to be detected in the source content set based on the weighted modal similarity. In this way, by determining the modal weight of each modality to weight the modal similarity of the content to be detected, the influence of the more important modality in different contents to be detected on the detection result is increased, and at the same time, the influence of the noise artificially introduced in the content to be detected on the accuracy of the detection result is reduced. Furthermore, the copyright information of the content to be detected is detected in the source content set according to the weighted modal similarity, so as to improve the accuracy of the content detection result and further improve the content detection efficiency. Description of the Drawings
[0044] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0045] Figure 1 It is a schematic diagram of an implementation scenario of a content detection method provided by an embodiment of the present application;
[0046] Figure 2 It is a schematic flowchart of a content detection method provided by an embodiment of the present application;
[0047] Figure 3 It is a schematic diagram of the overall structure of a content detection method provided by an embodiment of the present application;
[0048] Figure 4 It is another schematic flowchart of a content detection method provided by an embodiment of the present application;
[0049] Figure 5 It is a schematic diagram of the structure of a content detection device provided by an embodiment of the present application;
[0050] Figure 6 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0052] The embodiments of the present application provide a content detection method, device and computer-readable storage medium. Among them, the content detection device can be integrated in a computer device, and the computer device can be a server or a terminal device, etc.
[0053] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms. The terminal can be a device such as a tablet computer, a laptop computer, a desktop computer, a smart phone, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0054] Please refer to Figure 1 , taking the content detection device integrated in a computer device as an example, Figure 1 is a schematic diagram of an implementation scenario of the content detection method provided by an embodiment of the present application. Among them, the computer device can be a server or a terminal. The computer device can obtain the content to be detected and a set of source contents. The set of source contents includes at least one source content with copyright; perform multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and perform multi-modal feature extraction on the source content to obtain the source modal features of each modality; calculate the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality; determine the modal weight of each modality according to the to-be-detected modal features, and perform weighting on the modal similarity based on the modal weight; based on the weighted modal similarity, detect the copyright information of the content to be detected in the set of source contents.
[0055] It should be noted that Figure 1 the schematic diagram of the implementation environment scenario of the content detection method shown is only an example. The implementation environment scenario of the content detection method described in the embodiments of the present application is for more clearly explaining the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of content detection and the emergence of new business scenarios, the technical solutions provided by the present application are equally applicable to similar technical problems.
[0056] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0057] This embodiment will be described from the perspective of the content detection device. The content detection device can be specifically integrated in a computer device. The computer device can be a server, and this application does not make any restrictions here.
[0058] Please refer to Figure 2 , Figure 2It is a schematic flowchart of the content detection method provided by an embodiment of the present application. The content detection method includes:
[0059] In step 101, obtain the content to be detected and the source content set.
[0060] Among them, the content to be detected can be at least one content to be detected, which can be content such as video, text, image, audio, etc., or can be content including one or more types of the above types of content. The source content set can be an overall composed of at least one source content, and can include at least one source content with copyright. The source content can be at least one work with Intellectual Property (IP) copyright, which can be content such as video, text, image, audio, etc., or can be content including one or more types of the above types of content.
[0061] The content to be detected and the source content set can be multi-modal content with at least one modality. Among them, modality refers to the presentation method of information that can be received by people, such as images, text, audio, etc. Multi-modality can be application scenarios under multiple modalities. For example, short video scenarios (including image videos, voiceovers, text titles) can be multi-modal. In addition, the modality can also be a presentation method in other dimensions, such as pose features, etc., which are not limited here.
[0062] With the rapid development of Internet technology, more and more content such as videos, pictures, and texts is widely published on various platforms. Among them, there is a lot of content that infringes on the legitimate rights and interests of the rights holders in these published contents. These infringing contents avoid detection by infringement detection means by artificially introducing noise in the modality, which harms the legitimate rights and interests of the rights holders. For example, in the infringement detection scenario related to medical care, the entities with copyright may be various modalities such as medical popular science videos, pictures, texts, etc. In some medical classics, health and other products, the doctor popular science videos with copyright are often stolen by other platforms. At this time, it is necessary to combine algorithms and manual work to discover relevant infringing videos from other platforms to protect the legitimate rights and interests of the product. In these modalities, noise often exists in each collected modality. For example, in the infringement detection of videos, the collected infringing data may be processed through various blurring methods (such as filling in irrelevant content, reducing video clarity, etc.) to avoid infringement detection. The entity obtained after this processing will add noise through various artificial methods. This kind of noise is different from information loss and will make the modality information have a lower signal-to-noise ratio. Therefore, an information balance data processing method can be used to perform infringement detection on the content to be detected. Among them, information balance can refer to an equilibrium for information from different sources (i.e., modalities). For example, in a conversation (online meeting), the importance of the information in the audio modality can be greater than that in the video modality, and in a speech, the information in the audio modality and the information in text modalities such as speech scripts can be equally important.
[0063] Due to the strong heterogeneity of data between different modalities, such as text and image data, their processing methods and embedded spaces have large differences, and the information in a single modality also contains different degrees of noise. Information balance among various modalities is an extremely important but difficult task. In the existing technology, generally, a neural network model is simply used to adopt the attention mechanism to achieve information balance of multi-modal content. However, this method of the attention mechanism cannot well capture the importance between different modalities, nor can it well capture and process the noise information between modalities, resulting in a low accuracy of the content detection result, and further leading to a low content detection efficiency.
[0064] Therefore, the embodiments of the present application provide a content detection method. By designing an equilibrium mechanism to dynamically adapt to and learn the importance of each modality, the influence of the more important modality in different contents to be detected on the detection result can be increased, and at the same time, the influence of the artificially introduced noise in the content to be detected on the accuracy of the detection result can be reduced, so as to improve the accuracy of the content detection result and further improve the content detection efficiency. The content detection method provided by the embodiments of the present application will be described in detail below.
[0065] First, the content to be detected and the source content set can be obtained. The content to be detected can be the content for which infringement detection is to be carried out. The content to be detected can include at least one piece of content. The source content set can include at least one source content with copyright. Optionally, the source content can be obtained from the memory connected to the content detection device, or from other data storage terminals. It can also be obtained from the memory of an entity terminal, or from a virtual storage space such as a data set or a corpus.
[0066] In step 102, multi-modal feature extraction is performed on the content to be detected to obtain the to-be-detected modal features of each modality, and multi-modal feature extraction is performed on the source content to obtain the source modal features of each modality.
[0067] Among them, the to-be-detected modal features can be the feature information obtained by performing feature extraction on the content to be detected of each modality. This feature information can be the information used to characterize the features of the content to be detected of each modality. This feature information can be a feature vector. The feature vector can be a vector obtained by performing feature extraction on the content to be detected and vectorizing the extracted features. The specific vectorization method can use a feature extraction model to generate a vector corresponding to the feature according to the extracted feature, such as an embedding vector.
[0068] The source modal features can be the feature information obtained by performing feature extraction on the source content of each modality. This feature information can be the information used to characterize the features of the source content of each modality. This feature information can be a feature vector. The feature vector can be a vector obtained by performing feature extraction on the source content and vectorizing the extracted features. The specific vectorization method can use a feature extraction model to generate a vector corresponding to the feature according to the extracted feature, such as an embedding vector.
[0069] In this way, multi-modal feature extraction can be performed on the content to be detected to obtain the to-be-detected modal features of each modality, and multi-modal feature extraction can be performed on the source content to obtain the source modal features of each modality.
[0070] For example, multi-modal extraction can be performed on the content to be detected to obtain the to-be-detected modal data of each modality, and multi-modal extraction can be performed on the source content to obtain the source modal data of each modality; according to the modal type corresponding to each modality, the encoders corresponding to the to-be-detected modal data and the source modal data are respectively determined; based on the encoder corresponding to the to-be-detected modal data, feature extraction is performed on the to-be-detected modal data to obtain the to-be-detected modal features of each modality; based on the encoder corresponding to the source modal data, feature extraction is performed on the source modal data to obtain the source modal features of each modality.
[0071] Among them, the modal data to be detected can be the data corresponding to each modality in the content to be detected, and the source modal data can be the data corresponding to each modality in the source content. The modality type can be the type of each modality. For example, the modality type of the text modality is the text type, the modality type of the audio modality is the audio type, the modality type of the image modality is the image type, etc. A more appropriate encoder can be selected according to the modality type corresponding to each modality to further improve the accuracy of the detection result.
[0072] Specifically, the modal data in the content to be detected can be extracted to obtain the modal data to be detected for each modality, and the modal data in the source content can be extracted to obtain the source modal data for each modality. In order to obtain more accurate feature information, the encoders corresponding to the modal data to be detected and the source modal data can be determined respectively according to the modality type corresponding to each modality. For example, for the modal data of the image type, the encoder corresponding to the Residual Neural Network (ResNet) can be used, and for the modal data of the text type, the encoder corresponding to the Bidirectional Encoder Representations from Transformers (BERT) can be used, etc. Thus, based on the encoder corresponding to the modal data to be detected, the modal data to be detected can be feature-extracted to obtain the modal features to be detected for each modality, and based on the encoder corresponding to the source modal data, the source modal data can be feature-extracted to obtain the source modal features for each modality.
[0073] For example, please refer to Figure 3 , Figure 3 FIG.
[0074] In one embodiment, multi-modal feature extraction can be directly performed on the content to be detected to obtain the to-be-detected modal features of each modality, and multi-modal feature extraction can be directly performed on the source content to obtain the source modal features of each modality.
[0075] In step 103, calculate the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality.
[0076] Among them, the similarity can be the degree of similarity between two comparison objects, or can be described as the distance between two comparison objects. For example, the degree of similarity between two data, the degree of similarity between two vectors. Here, it can be the degree of similarity between the to-be-detected modal features and the source modal features of the corresponding modality. The modal similarity can be the similarity between each to-be-detected modal feature and the source modal feature of the corresponding modality.
[0077] In order to obtain the similarity between the content to be detected and the source content to measure whether the content to be detected causes infringement of the source content, the similarity between the to-be-detected modal features and the source modal features of the corresponding modality can be calculated to obtain the modal similarity of each modality.
[0078] For example, please continue to refer to Figure 3 , assuming that the content to be detected includes the to-be-detected modal features of three modalities: text modal features, image modal features, and audio modal features, and the source content includes the source modal features of three modalities: text modal features, image modal features, and audio modal features. Then, the similarity between the to-be-detected modal features of each modality and the source modal features of the corresponding modality can be calculated respectively. For example, the similarity between the text modal features corresponding to the content to be detected and the text modal features corresponding to the source content can be calculated to obtain the modal similarity 1 of the text modality. The similarity between the image modal features corresponding to the content to be detected and the image modal features corresponding to the source content can be calculated to obtain the modal similarity 2 of the image modality. The similarity between the audio modal features corresponding to the content to be detected and the audio modal features corresponding to the source content can be calculated to obtain the modal similarity 3 of the audio modality.
[0079] In step 104, determine the modal weight of each modality according to the to-be-detected modal features, and weight the modal similarity based on the modal weight.
[0080] Among them, the modality weight can be the weight value of each modality in the content to be detected, which can represent the importance of each modality in the content to be detected. For example, in song content, the audio modality is generally the most important, so the modality weight corresponding to the audio modality is the largest. In article content, the text modality is generally the most important weight, and the modality weight corresponding to the text modality is the largest. In addition, by weighting the modality weights of each modality, the influence of unimportant modalities on the content detection result can be weakened. At the same time, the noise introduced in a certain unimportant modality can be further reduced from affecting the accuracy of the detection result, so as to achieve noise reduction for unimportant modalities and improve the accuracy of content detection.
[0081] Thus, based on the modality features to be detected, the modality weight of each modality in the content to be detected can be determined, and the modality similarity of each modality can be weighted based on the modality weight to obtain the weighted modality similarity, that is, the similarity between the content to be detected and the source content.
[0082] For example, please continue to refer to Figure 3 , assuming that the content to be detected includes the modality features to be detected of three modalities: text modality feature, image modality feature, and audio modality feature. The modality weight of the text modality in the content to be detected can be determined as modality weight 1 based on the text modality feature corresponding to the content to be detected. The modality weight of the image modality in the content to be detected can be determined as modality weight 2 based on the image modality feature corresponding to the content to be detected. The modality weight of the audio modality in the content to be detected can be determined as modality weight 3 based on the audio modality feature corresponding to the content to be detected. Thus, the modality similarity can be weighted based on modality weight 1, modality weight 2, and modality weight 3 to obtain the weighted modality similarity. For example, assuming that the weighted modality similarity is S, and modality weight 1, modality weight 2, and modality weight 3 are w1, w2, and w3 respectively, the modality similarity corresponding to the text modality is S1, the modality similarity corresponding to the image modality is S2, and the modality similarity corresponding to the audio modality is S3. Then the weighted modality similarity corresponding to the text modality can be w1×S1, the weighted modality similarity corresponding to the image modality can be w2×S2, and the weighted modality similarity corresponding to the audio modality can be w3×S3.
[0083] Among them, there are various ways to determine the modality weight of each modality according to the modality features to be detected. For example, the modality importance score corresponding to the modality features to be detected can be identified; the modality quality score of the modality features to be detected can be detected; the modality importance score and the modality quality score of the corresponding modality can be fused to obtain the modality weight of each modality.
[0084] Among them, the modal importance score can be the weight value of each modality in the content to be detected, which can represent the importance degree of each modality in the content to be detected. For example, in song content, the audio modality is generally the most important, so the modal importance score corresponding to the audio modality is the largest. In article content, the text modality is generally the most important weight, and the modal importance score corresponding to the text modality is the largest. The modal quality score can be the quality score of each modality in the content to be detected, which can represent the magnitude of noise in each modality of the content to be detected. If the noise is large, the quality score is low; if the noise is small, the quality score is high. For example, the noise detection can be performed on the data of each modality in the content to be detected, and then the modal quality score of each modality can be determined according to the noise detection result. The noise detection can also be performed on the modality features to be detected in each modality of the content to be detected, and then the modal quality score of each modality can be determined according to the noise detection result.
[0085] Specifically, the modal importance score corresponding to the modality features to be detected in each modality of the content to be detected can be identified, the modal quality score of the modality features to be detected can be detected, and then the modal importance score and the modal quality score of the corresponding modality can be fused to obtain the modal weight of each modality. For example, the modal importance score and the modal quality score of the corresponding modality can be weighted to obtain the modal weight of each modality.
[0086] Optionally, according to the modality features to be detected, the trained content detection model can be used to determine the modal weight of each modality, and the modal similarity can be weighted based on the modal weight. Among them, the trained content detection model can be a model for detecting content after training, which can be used to detect the content to be detected to obtain the copyright information corresponding to the content to be detected.
[0087] For example, the trained content detection model can obtain at least one pair of content samples, and use the preset content detection model to predict the modal weight of each modality based on the detected content sample and the source content sample to obtain the predicted modal weight. Based on the predicted modal weight, the modal similarity of the content sample pair can be predicted to obtain the predicted modal similarity of each modality. According to the predicted modal similarity and the labeled copyright information, the preset content detection model can be converged to obtain the trained content detection model. Specifically, it can be as follows:
[0088] (1) Obtain at least one pair of content samples.
[0089] Among them, the content sample pair may include a detected content sample, a source content sample, and annotated copyright information. There is a corresponding relationship among the detected content sample, the source content sample, and the annotated copyright information in each content sample pair. The annotated copyright information in each content sample pair is the corresponding annotated copyright information between the detected content sample and the source content sample therein. The detected content sample may be at least one content sample for detection, which may be content samples such as videos, texts, images, audios, etc., or may be content samples including one or more of the above types of content. The source content sample may be at least one content sample with IP copyright, which may be content samples such as videos, texts, images, audios, etc., or may be content samples including one or more of the above types of content.
[0090] The annotated copyright information may be information annotating whether each detected content sample is infringed, or may be information on whether there is infringement between content sample pairs. For example, 1 may indicate infringement, and 0 may indicate non-infringement. If a content sample pair is an infringing sample, that is, the detected content sample in the content sample pair infringes the source content sample in the content sample pair, the annotated copyright information of the detected content sample may be marked as 1. If a detected content sample is a non-infringing sample, the annotated copyright information of the detected content sample may be marked as 0, etc.
[0091] In this way, at least one content sample pair can be obtained, and training can be carried out through the content sample pair to obtain a trained content detection model.
[0092] (2) Using a preset content detection model, predict the modal weight of each modality based on the detected content sample and the annotated copyright information to obtain a predicted modal weight.
[0093] Among them, the predicted modal weight may be the modal weight of each modality obtained by predicting through a preset content detection model according to the detected content sample and the source content sample.
[0094] In this way, a preset content detection model can be used to predict the modal weight of each modality based on the detected content sample and the source content sample to obtain a predicted modal weight. For example, the detected content sample and the source content sample can be input into the preset content detection model for prediction to obtain the predicted modal weight corresponding to each modality.
[0095] Optionally, the preset content detection model may include a weight determination sub-model. The step of using the preset content detection model to predict the modal weight of each modality based on the detected content sample and the source content sample to obtain a predicted modal weight may include:
[0096] (2.1) Use the weight determination sub-model to perform multi-modal feature extraction on the detected content sample to obtain the sample modal features of each modality.
[0097] Among them, the weight determination sub-model can be a model for determining the modal weight of each modality, and the sample modal feature can be the modal feature of each modality corresponding to the content sample to be detected. For example, it can be the embedding corresponding to each modality in the content sample to be detected. In this way, the weight determination sub-model can be used to perform multi-modal feature extraction on the detected content sample to obtain the sample modal features of each modality. Optionally, a preset content detection model can be used to perform multi-modal feature extraction on the detected content sample to obtain the sample modal features of each modality.
[0098] (2.2) Calculate the correlation coefficient between the sample modal features of each modality and the marked copyright information.
[0099] Among them, the correlation coefficient can be the correlation between the sample modal features of each modality and the corresponding copyright information. In this way, the correlation coefficient between the sample modal features of each modality and the marked copyright information can be calculated. For example, the point-wise mutual information function can be used to measure the correlation between the sample modal features of each modality and the marked copyright information, so as to obtain the correlation coefficient between the sample modal features of each modality and the marked copyright information.
[0100] (2.3) Based on the correlation coefficient, determine the predicted modal weight of each modality.
[0101] In this way, based on the correlation coefficient between the sample modal features of each modality and the marked copyright information, the predicted modal weight of each modality can be determined. For example, assume that the marked copyright information can be C, the detected content sample is Q, the source content sample is P, the detected content sample can have M modalities, and at the same time, assume that the sample modal feature corresponding to modality i can be R i , and the marked copyright information corresponding to modality i can be C, then the weight determination sub-model can be expressed as
[0102] ω i =sofimax(I(C, R i ))
[0103] Among them, ω i can also be the predicted modal weight corresponding to modality i, and softmax() is the normalized exponential function (softmax function), which is a differentiable and smooth function. Optionally, a temperature parameter can be added to converge the softmax function. I(C, R i ) can represent the correlation function. Optionally, since C and R iBoth are discrete feature vector representations. Therefore, the point mutual information function can be used as a metric to measure C and R i The correlation between them.
[0104] Optionally, when the detected content sample and the source content sample are a set including multiple sub-content samples. For example, assuming there are multiple detected sub-content samples q in the detected content sample Q and multiple source sub-content samples p in the source content sample P, and each sub-content sample can include M modalities, then for each q ∈ Q (q belongs to Q), there will be |P| values of C(p, q). Therefore, the mean sum can be calculated for P to obtain C i = sum(C(p, q)) / |P|, where sum(C(p, q)) can be the sum obtained by accumulating the labeled copyright information between each detected sub-content sample q and each source content sample P in the i-th modality, and R i can be the set of sample modality features of each q in the detected content sample Q in modality i.
[0105] (3) Based on the predicted modality weight, predict the modality similarity of the content sample pair to obtain the predicted modality similarity of each modality.
[0106] Among them, the predicted modality similarity can be the similarity between the detected content sample corresponding to each modality and the corresponding source content sample obtained based on the predicted modality weight. Thus, based on the predicted modality weight, the modality similarity of the content sample pair can be predicted to obtain the predicted modality similarity of each modality. For example, the predicted modality similarity for modality i can be expressed as ω i *s i , where ω i can be the predicted modality weight corresponding to modality i, and s i can be the predicted modality similarity of the content sample pair corresponding to modality i. For example, it can be expressed as
[0107] s i = ENC i (P) * ENC i (Q)
[0108] Among them, ENC i (Q) represents the modality feature corresponding to the detected content sample in modality i, and ENC i (P) represents the modality feature corresponding to the source content sample in modality i, and ENC i() can represent the encoder corresponding to modality i. To more accurately extract data for each modality, different encoders can be selected for feature extraction of data of different modality types. For example, for data of the image modality, a pre-trained ResNet can be used for feature extraction, and for data of the text modality, a pre-trained BERT can be used for feature extraction. Q can be the data of the detected content sample input into the encoder corresponding to modality i. For example, when modality i is the text modality, Q can be text data. P can be the data of the source content sample input into the encoder corresponding to modality i. For example, when modality i is the image modality, P can be image data.
[0109] (4) According to the predicted modality similarity and the labeled copyright information, converge the preset content detection model to obtain a trained content detection model.
[0110] In this way, according to the modality similarity and the labeled copyright information, the preset content detection model can be converged to obtain a trained content detection model.
[0111] For example, according to the predicted modality similarity of each modality, the total predicted modality similarity of the content sample pair can be calculated; based on the total predicted modality similarity and the corresponding labeled copyright information, the target loss information of each content sample pair can be determined; based on the target loss information, converge the preset content detection model to obtain a trained content detection model.
[0112] Among them, the total predicted modality similarity can be the sum obtained by accumulating the predicted modality similarities of each modality corresponding to the content sample pair, and the target loss information can be the loss between the total predicted modality similarity corresponding to the content sample pair and the labeled copyright information.
[0113] Specifically, according to the predicted modality similarity of each modality, the total predicted modality similarity of the content sample pair can be calculated. For example, the predicted modality similarities of each modality can be accumulated to obtain the total predicted modality similarity of the content sample pair. The calculation formula of the total predicted modality similarity can be
[0114]
[0115] Among them, ∑ (Sigma) represents the summation symbol in mathematics, which is mainly used to find the sum of multiple numbers. Here, it can represent the summation of the predicted modality similarities ω i *s i for M modalities, and s i can be the predicted modality similarity of the content sample pair corresponding to modality i.
[0116] Thus, based on the total predicted modal similarity and the corresponding labeled copyright information, the target loss information for each content sample pair can be determined. For example, based on the total predicted modal similarity and the labeled copyright information, a loss function for converging the preset content detection model can be determined. For example, the loss function can be expressed as
[0117] L = ∑|C - S s |
[0118] where S s represents the total predicted modal similarity of the content sample pair, that is, the total predicted modal similarity between the content sample Q and the source content sample P in the content sample pair. C can represent the labeled copyright information in the content sample pair. Thus, the target loss information for each content sample pair can be calculated through this loss function, and based on this target loss information, the preset content detection model can be converged. When the convergence condition is met, the trained content detection model can be obtained.
[0119] In this way, based on the trained content detection model, the modal weight of each modality in the content to be detected can be obtained. Furthermore, based on the modal weight of each modality, the content to be detected can be detected to more accurately determine whether there is an infringement behavior in the content to be detected. In addition, based on the modal weight of each modality, an encoder that is more matched to each modality can be selected, thereby further improving the accuracy of content detection.
[0120] In step 105, based on the weighted modal similarity, the copyright information of the content to be detected is detected in the source content set.
[0121] Among them, the weighted modal similarity can be the similarity obtained by weighting the modal similarity of the corresponding modality based on the modal weight of each modality. For example, assume that the modal weight corresponding to modality l is w l , and the modal similarity corresponding to modality l is s l , then the weighted modal similarity can be w l ×s l .
[0122] The copyright information can be information indicating whether the content to be detected is infringing. Based on the weighted modal similarity, it can be detected in the source content set, and the copyright information of the content to be detected can be determined according to the detection result. For example, when a target source content is detected in the source content set based on the weighted modal similarity, it can indicate that the content to be detected is an infringing content, and the object of infringement is the target source content. When no target source content is detected in the source content set based on the weighted modal similarity, it can indicate that there is no infringing content in the content to be detected in the source content set, that is, the content to be detected does not infringe on other works.
[0123] For example, it is possible to determine whether there is a potential for infringement of the content to be detected by calculating the similarity between the content to be detected and the source content in each modality. A high similarity can indicate a high similarity between the content to be detected and the corresponding source content, that is, it can be determined that the content to be detected is an infringing content.
[0124] Optionally, the weighted modality similarities corresponding to each modality can be fused to obtain the total weighted modality similarity corresponding to each source content; based on this total weighted modality similarity, target source content can be screened out from the source content set; based on this target source content, the copyright information of the content to be detected can be determined.
[0125] Among them, the total weighted modality similarity can be a value obtained by accumulating the weighted modality similarities between each source content and the content to be detected in each modality. For example, assume that there are 3 modalities in the content to be detected, namely Modality 1, Modality 2, and Modality 3. Among them, the modality weight corresponding to Modality 1 is w1, and the corresponding modality similarity is s1, then the weighted modality similarity corresponding to Modality 1 can be w1×s1, the modality weight corresponding to Modality 2 is w2, and the corresponding modality similarity is s2, then the weighted modality similarity corresponding to Modality 2 can be w2×s2, the modality weight corresponding to Modality 3 is w3, and the corresponding modality similarity is s3, then the weighted modality similarity corresponding to Modality 3 can be w3×s3, then the total weighted modality similarity can be S_total = w1×s1 + w2×s2 + w3×s3.
[0126] The target source content can be at least one source content screened out from the source content set based on the total weighted modality similarity. For example, a similarity threshold can be preset. The similarity threshold can be a critical value of similarity. When the total weighted modality similarity is greater than the similarity threshold, it can be considered that the content to be detected infringes the source content corresponding to the total weighted modality similarity, then the source content corresponding to the total weighted modality similarity can be determined as the target source content. When the total weighted modality similarity is not greater than the similarity threshold, it can be determined that the content to be detected does not infringe the source content in the source content set.
[0127] Specifically, the weighted post-modal similarity corresponding to each modality can be fused to obtain the total weighted post-modal similarity corresponding to each source content. For example, the weighted post-modal similarities corresponding to each modality can be accumulated to obtain the total weighted post-modal similarity corresponding to each source content. Furthermore, based on this total weighted post-modal similarity, the target source content can be screened out from the source content set. For example, a similarity threshold can be preset, and at least one source content with a total weighted post-modal similarity greater than this similarity threshold can be screened out from the source content set to obtain the target source content, so that the copyright information of the content to be detected can be determined based on this target source content.
[0128] For example, assume that there are source content 1, source content 2, and source content 3 in the source content set. Among them, the total weighted post-modal similarity corresponding to source content 1 is 2.1, the total weighted post-modal similarity corresponding to source content 2 is 0.4, and the total weighted post-modal similarity corresponding to source content 3 is 5.2. At this time, if the preset similarity threshold is 5, then based on the total weighted post-modal similarity corresponding to each source content, the target source content greater than this similarity threshold of 5 screened out from the source content set is source content 3, that is, it can be obtained that the copyright information may be that the content to be detected infringes on source content 3. When the preset similarity threshold is 6, based on the total weighted post-modal similarity corresponding to each source content, no target source content greater than this similarity threshold of 6 can be screened out from the source content set, that is, it can be obtained that the copyright information is that the content to be detected does not infringe on source content 1, source content 2, and source content 3. When the preset similarity threshold is 2, then based on the total weighted post-modal similarity corresponding to each source content, the target source content greater than this similarity threshold of 2 screened out from the source content set is source content 1 and source content 3, that is, it can be obtained that the copyright information may be that the content to be detected infringes on source content 1 and source content 3.
[0129] As can be seen from the above, in the embodiment of the present application, by obtaining the content to be detected and the source content set, the source content set includes at least one source content with copyright; performing multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and performing multi-modal feature extraction on the source content to obtain the source modal features of each modality; calculating the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality; determining the modal weight of each modality according to the to-be-detected modal features, and weighting the modal similarity based on the modal weight; detecting the copyright information of the content to be detected in the source content set based on the weighted modal similarity. In this way, by determining the modal weight of each modality to weight the modal similarity of the content to be detected, the influence of the more important modality in different contents to be detected on the detection result is increased, and at the same time, the influence of the noise artificially introduced in the content to be detected on the accuracy of the detection result is reduced. Furthermore, the copyright information of the content to be detected is detected in the source content set according to the weighted modal similarity, thereby improving the accuracy of the content detection result and further improving the content detection efficiency.
[0130] According to the method described in the above embodiment, the following will be further described in detail by way of examples.
[0131] In this embodiment, it will be described by taking the content detection device being specifically integrated in a computer device as an example. Among them, the content detection method is specifically described with the server as the execution subject.
[0132] For a better description of the embodiment of the present application, please refer to Figure 4 . As Figure 4 shown, Figure 4 is another schematic flowchart of the content detection method provided by the embodiment of the present application. The specific process is as follows:
[0133] In step 201, the server obtains at least one pair of content samples, performs multi-modal feature extraction on the content sample to be detected by using the weight determination sub-model to obtain the sample modal features of each modality, calculates the correlation coefficient between the sample modal features of each modality and the annotated copyright information, and determines the predicted modal weight of each modality based on the correlation coefficient.
[0134] Among them, the server can obtain at least one content sample pair, and then can use the weight determination sub-model to perform multi-modal feature extraction on the detected content sample to obtain the sample modal features of each modality. Thus, the point mutual information function can be used to measure the correlation between the sample modal features of each modality and the labeled copyright information, so as to obtain the correlation coefficient between the sample modal features of each modality and the labeled copyright information. Based on the correlation coefficient between the sample modal features of each modality and the labeled copyright information, the predicted modal weight of each modality can be determined. For example, assume that the labeled copyright information can be C, the detected content sample is Q, the source content sample is P, the detected content sample can have M modalities, and at the same time, assume that the sample modal feature corresponding to modality i can be R i , the labeled copyright information corresponding to modality i can be C, then the weight determination sub-model can be expressed as
[0135] ω i =sofimax(I(C, R i ))
[0136] Among them, ω i can also be the predicted modal weight corresponding to modality i, softmax() is the normalized exponential function (softmax function), which is a differentiable and smooth function. Optionally, a temperature parameter can be added to converge the softmax function. I(C, R i ) can represent the correlation function. Optionally, since both C and R i are represented by discrete feature vectors, the point mutual information function can be used as a measure to measure the correlation between C and R i .
[0137] Optionally, when the detected content sample and the source content sample are a set including multiple sub-content samples, for example, assume that there are multiple detected sub-content samples q in the detected content sample Q, and multiple source sub-content samples p in the source content sample P, and each sub-content sample can include M modalities. Then, for each q ∈ Q (q belongs to Q), there will be |P| values of C(p,q). Therefore, the server can perform mean summation on P to obtain C i =sum(C(p,q)) / |P|, where sum(C(p,q)) can be the sum obtained by accumulating the labeled copyright information between each detected sub-content sample q in modality i and each source content sample P, and R i can be the set of sample modal features of each q in the detected content sample Q in modality i.
[0138] In step 202, the server predicts the modal similarity of the content sample pair based on the predicted modal weights, obtains the predicted modal similarity of each modality, and calculates the total predicted modal similarity of the content sample pair according to the predicted modal similarity of each modality.
[0139] For example, the server can predict the modal similarity of the content sample pair based on the predicted modal weights, obtain the predicted modal similarity of each modality. For example, the predicted modal similarity for modality i can be expressed as ω i *s i , where ω i can be the predicted modal weight corresponding to modality i, and s i can be the predicted modal similarity of the content sample pair corresponding to modality i. For example, it can be expressed as
[0140] s i =ENC i (P)*ENC i (Q)
[0141] where ENC i (Q) represents the modal feature corresponding to the detected content sample in modality i, and ENC i (P) represents the modal feature corresponding to the source content sample in modality i. ENC i () can represent the encoder corresponding to modality i. To more accurately extract data for each modality, different encoders can be selected for data of different modality types for feature extraction. For example, for image modality data, pre-trained ResNet can be used for feature extraction, and for text modality data, pre-trained BERT can be used for feature extraction. Q can be the data input by the detected content sample into the corresponding encoder of modality i. For example, when modality i is the text modality, Q can be text data. P can be the data input by the source content sample into the corresponding encoder of modality i. For example, when modality i is the image modality, P can be image data.
[0142] After obtaining the predicted modal similarity of each modality, the server can accumulate the predicted modal similarity of each modality to obtain the total predicted modal similarity of the content sample pair. The calculation formula for the total predicted modal similarity can be
[0143]
[0144] where ∑ (Sigma) represents the summation symbol in mathematics, mainly used to find the sum of multiple numbers. Here, it can represent the summation of the predicted modal similarities ω i *s i corresponding to M modalities, and s iIt can be the predicted modal similarity corresponding to the content sample pair in modality i.
[0145] In step 203, the server determines the target loss information for each content sample pair based on the total predicted modal similarity and the corresponding labeled copyright information, and converges the preset content detection model based on this target loss information to obtain a trained content detection model.
[0146] For example, the server can determine the target loss information for each content sample pair according to the total predicted modal similarity and the corresponding labeled copyright information. For example, it can determine the loss function for converging the preset content detection model according to the total predicted modal similarity and the labeled copyright information. For example, the loss function can be expressed as
[0147] L=∑|C - S s |
[0148] where S s represents the total predicted modal similarity of the content sample pair, that is, the total predicted modal similarity between the content sample Q and the source content sample P in the content sample pair. C can represent the labeled copyright information in the content sample pair. Thus, the target loss information for each content sample pair can be calculated through this loss function, and the preset content detection model can be converged based on this target loss information. When the convergence condition is met, a trained content detection model can be obtained.
[0149] In step 204, the server obtains the content to be detected and the source content set, performs multi-modal extraction on the content to be detected to obtain the to-be-detected modal data for each modality, and performs multi-modal extraction on the source content to obtain the source modal data for each modality.
[0150] Among them, after obtaining the content to be detected and the source content set, the server can perform multi-modal extraction on the content to be detected to obtain the to-be-detected modal data for each modality, and perform multi-modal extraction on the source content to obtain the source modal data for each modality. For example, the server can extract the modal data in the content to be detected to obtain the to-be-detected modal data for each modality, and can extract the modal data in the source content to obtain the source modal data for each modality.
[0151] In step 205, the server respectively determines the encoders corresponding to the to-be-detected modal data and the source modal data according to the modal type corresponding to each modality, extracts features from the to-be-detected modal data based on the encoder corresponding to the to-be-detected modal data to obtain the to-be-detected modal features for each modality, and extracts features from the source modal data based on the encoder corresponding to the source modal data to obtain the source modal features for each modality.
[0152] For example, in order to obtain more accurate feature information, the server can determine the encoders corresponding to the to-be-detected modal data and the source modal data respectively according to the modal type corresponding to each modality. For example, for modal data of the image type, the encoder corresponding to the Residual Neural Network (ResNet) can be used, and for modal data of the text type, the encoder corresponding to the Bidirectional Encoder Representations from Transformers (BERT) can be used, etc. Thus, based on the encoder corresponding to the to-be-detected modal data, feature extraction can be performed on the to-be-detected modal data to obtain the to-be-detected modal features of each modality, and based on the encoder corresponding to the source modal data, feature extraction can be performed on the source modal data to obtain the source modal features of each modality.
[0153] For example, please continue to refer to Figure 3 and it can be assumed that there is data of modalities such as text, image, audio, etc. in the to-be-detected content and the source content. The encoders corresponding to the to-be-detected modal data and the source modal data are determined according to the modal type of each modality. For example, for the to-be-detected modal data and the source modal data with the modal type of text type, it can be determined that the corresponding encoder is Encoder 1, for the to-be-detected modal data and the source modal data with the modal type of image type, it can be determined that the corresponding encoder is Encoder 2, and for the to-be-detected modal data and the source modal data with the modal type of audio type, it can be determined that the corresponding encoder is Encoder 3. Thus, feature extraction can be performed on the to-be-detected modal data of each modality according to Encoder 1, Encoder 2, and Encoder 3, etc., to obtain the to-be-detected modal features such as the text modal feature, the image modal feature, and the audio modal feature corresponding to the to-be-detected content, and feature extraction can be performed on the source modal data of each modality according to Encoder 1, Encoder 2, and Encoder 3, etc., to obtain the source modal features such as the text modal feature, the image modal feature, and the audio modal feature corresponding to the source content.
[0154] In one embodiment, the server can directly perform multi-modal feature extraction on the to-be-detected content to obtain the to-be-detected modal features of each modality, and can directly perform multi-modal feature extraction on the source content to obtain the source modal features of each modality.
[0155] In step 206, the server calculates the similarity between the to-be-detected modal feature and the source modal feature of the corresponding modality to obtain the modal similarity of each modality, determines the modal weight of each modality by using the trained content detection model according to the to-be-detected modal feature, and weights the modal similarity based on the modal weight.
[0156] For example, in order to obtain the similarity between the content to be detected and the source content to measure whether the content to be detected infringes on the source content, the server can calculate the similarity between the modality features to be detected and the source modality features of the corresponding modality, obtain the modality similarity of each modality, and then determine the modality weight of each modality using the trained content detection model based on the modality features to be detected, and weight the modality similarity based on the modality weight. Among them, the trained content detection model can be a model trained to detect content, which can be used to detect the content to be detected to obtain the copyright information corresponding to the content to be detected.
[0157] For example, please continue to refer to Figure 3 , assuming that the content to be detected includes modality features to be detected in three modalities: text modality features, image modality features, and audio modality features, the server can determine that the modality weight of the text modality in the content to be detected is modality weight 1 based on the text modality features corresponding to the content to be detected, can determine that the modality weight of the image modality in the content to be detected is modality weight 2 based on the image modality features corresponding to the content to be detected, and can determine that the modality weight of the audio modality in the content to be detected is modality weight 3 based on the audio modality features corresponding to the content to be detected. Thus, the modality similarity can be weighted based on modality weight 1, modality weight 2, and modality weight 3 to obtain the weighted modality similarity. For example, assuming that the weighted modality similarity is S, modality weight 1, modality weight 2, and modality weight 3 are w1, w2, and w3 respectively, the modality similarity corresponding to the text modality is S1, the modality similarity corresponding to the image modality is S2, and the modality similarity corresponding to the audio modality is S3. Then, the weighted modality similarity corresponding to the text modality can be w1×S1, the weighted modality similarity corresponding to the image modality can be w2×S2, and the weighted modality similarity corresponding to the audio modality can be w3×S3.
[0158] In step 207, the server fuses the weighted modality similarities corresponding to each modality to obtain the total weighted modality similarity corresponding to each source content, screens out the target source content from the source content set according to the total weighted modality similarity, and determines the copyright information of the content to be detected based on the target source content.
[0159] Among them, the server can fuse the weighted modal similarities corresponding to each modality to obtain the total weighted modal similarity corresponding to each source content. For example, the weighted modal similarities corresponding to each modality can be accumulated to obtain the total weighted modal similarity corresponding to each source content. Furthermore, based on this total weighted modal similarity, the target source content can be screened out from the source content set. For example, a similarity threshold can be preset, and at least one source content with a total weighted modal similarity greater than this similarity threshold can be screened out from the source content set to obtain the target source content. Thus, based on this target source content, the copyright information of the content to be detected can be determined.
[0160] For example, assume that there are source content 1, source content 2, and source content 3 in the source content set. Among them, the total weighted modal similarity corresponding to source content 1 is 2.1, the total weighted modal similarity corresponding to source content 2 is 0.4, and the total weighted modal similarity corresponding to source content 3 is 5.2. At this time, if the preset similarity threshold is 5, the server can screen out the target source content greater than this similarity threshold 5 from the source content set, which is source content 3. That is, it can be obtained that the copyright information may be that the content to be detected infringes on source content 3. When the preset similarity threshold is 6, according to the total weighted modal similarity corresponding to each source content, no target source content greater than this similarity threshold 6 can be screened out from the source content set. That is, it can be obtained that the copyright information is that the content to be detected does not infringe on source content 1, source content 2, and source content 3. When the preset similarity threshold is 2, then according to the total weighted modal similarity corresponding to each source content, the target source content greater than this similarity threshold 2 screened out from the source content set is source content 1 and source content 3. That is, it can be obtained that the copyright information may be that the content to be detected infringes on source content 1 and source content 3.
[0161] As can be seen from the above, in the embodiment of the present application, the server obtains at least one content sample pair, uses the weight determination sub-model to perform multi-modal feature extraction on the detected content sample to obtain the sample modal features of each modality, calculates the correlation coefficient between the sample modal features of each modality and the annotated copyright information, and determines the predicted modal weight of each modality based on the correlation coefficient; the server predicts the modal similarity of the content sample pair based on the predicted modal weight, obtains the predicted modal similarity of each modality, and determines the similarity loss information corresponding to each modality according to the predicted modal similarity and the annotated copyright information; the server fuses the similarity loss information to obtain the target loss information of the content sample pair, and converges the preset content detection model based on the target loss information to obtain a trained content detection model; the server obtains the content to be detected and the source content set, performs multi-modal extraction on the content to be detected to obtain the detected modal data of each modality, and performs multi-modal extraction on the source content to obtain the source modal data of each modality; the server respectively determines the encoders corresponding to the detected modal data and the source modal data according to the modal type corresponding to each modality, performs feature extraction on the detected modal data based on the encoder corresponding to the detected modal data to obtain the detected modal features of each modality, and performs feature extraction on the source modal data based on the encoder corresponding to the source modal data to obtain the source modal features of each modality; the server calculates the similarity between the detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality, determines the modal weight of each modality according to the detected modal features by using the trained content detection model, and weights the modal similarity based on the modal weight; the server fuses the weighted modal similarities corresponding to each modality to obtain the total weighted modal similarity corresponding to each source content, screens out the target source content from the source content set according to the total weighted modal similarity, and determines the copyright information of the content to be detected based on the target source content. In this way, by determining the predicted modal similarity based on the correlation coefficient between the detected content sample and the annotated copyright information, and determining the similarity loss information corresponding to each modality according to the predicted modal similarity and the annotated copyright information, the preset content detection model is trained to obtain a trained content detection model. Furthermore, in the application process, the modal weight of each modality can be determined to weight the modal similarity of the content to be detected, so as to increase the influence of the more important modality in different contents to be detected on the detection result, and at the same time reduce the influence of the noise introduced artificially in the content to be detected on the accuracy of the detection result. Then, according to the weighted modal similarity, the copyright information of the content to be detected is detected in the source content set, thereby improving the accuracy of the content detection result and further improving the content detection efficiency.
[0162] To better implement the above method, an embodiment of the present invention further provides a content detection device, which can be integrated in a computer device, and the computer device can be a server.
[0163] For example, as Figure 5 shown, it is a schematic structural diagram of the content detection device provided by an embodiment of the present application. The content detection device may include an acquisition unit 301, a feature extraction unit 302, a calculation unit 303, a determination unit 304, and a detection unit 305, as follows:
[0164] The acquisition unit 301 is configured to acquire the content to be detected and a set of source contents, and the set of source contents includes at least one source content with copyright;
[0165] The feature extraction unit 302 is configured to perform multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and perform multi-modal feature extraction on the source content to obtain the source modal features of each modality;
[0166] The calculation unit 303 is configured to calculate the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality;
[0167] The determination unit 304 is configured to determine the modal weight of each modality according to the to-be-detected modal features, and weight the modal similarity based on the modal weight;
[0168] The detection unit 305 is configured to detect the copyright information of the content to be detected in the set of source contents based on the weighted modal similarity.
[0169] In one embodiment, the determination unit 304 includes:
[0170] A modal importance score recognition sub-unit, configured to recognize the modal importance score corresponding to the to-be-detected modal features;
[0171] A modal quality score detection sub-unit, configured to detect the modal quality score of the to-be-detected modal features;
[0172] A coefficient fusion sub-unit, configured to fuse the modal importance score and the modal quality score of the corresponding modality to obtain the modal weight of each modality.
[0173] In one embodiment, the detection unit 305 includes:
[0174] A similarity fusion sub-unit, configured to fuse the weighted modal similarities corresponding to each modality to obtain the total weighted modal similarity corresponding to each source content;
[0175] A screening subunit, configured to screen out target source content from the source content set according to the total weighted modal similarity;
[0176] A copyright information determination subunit, configured to determine the copyright information of the content to be detected based on the target source content.
[0177] In one embodiment, the feature extraction unit 302 includes:
[0178] A modal extraction subunit, configured to perform multi-modal extraction on the content to be detected to obtain the to-be-detected modal data of each modality, and perform multi-modal extraction on the source content to obtain the source modal data of each modality;
[0179] An encoder determination subunit, configured to respectively determine the encoders corresponding to the to-be-detected modal data and the source modal data according to the modal type corresponding to each modality;
[0180] A to-be-detected extraction subunit, configured to perform feature extraction on the to-be-detected modal data based on the encoder corresponding to the to-be-detected modal data to obtain the to-be-detected modal features of each modality;
[0181] A source extraction subunit, configured to perform feature extraction on the source modal data based on the encoder corresponding to the source modal data to obtain the source modal features of each modality.
[0182] In one embodiment, the determination unit 304 includes:
[0183] A model determination subunit, configured to determine the modal weight of each modality by using a trained content detection model according to the to-be-detected modal features, and weight the modal similarity based on the modal weight.
[0184] In one embodiment, the content detection device further includes:
[0185] A sample acquisition unit, configured to acquire at least one content sample pair, where the content sample pair includes a detection content sample, a source content sample, and an annotated copyright information;
[0186] A weight prediction unit, configured to predict the modal weight of each modality by using a preset content detection model based on the detection content sample and the annotated copyright information to obtain a predicted modal weight;
[0187] A similarity prediction unit, configured to predict the modal similarity of the content sample pair based on the predicted modal weight to obtain the predicted modal similarity of each modality;
[0188] A convergence unit, configured to converge the preset content detection model according to the predicted modal similarity and the annotated copyright information to obtain a trained content detection model.
[0189] In one embodiment, the convergence unit includes:
[0190] A calculation subunit, configured to calculate the total predicted modal similarity of the content sample pair according to the predicted modal similarity of each modality.
[0191] A loss information determination subunit, configured to determine the target loss information of each content sample pair based on the total predicted modal similarity and the corresponding labeled copyright information.
[0192] A convergence subunit, configured to converge the preset content detection model based on the target loss information to obtain a trained content detection model.
[0193] In one embodiment, the weight prediction unit includes:
[0194] A sample feature extraction subunit, configured to perform multi-modal feature extraction on the detected content sample by using the weight determination sub-model to obtain the sample modal features of each modality.
[0195] A correlation coefficient calculation subunit, configured to calculate the correlation coefficient between the sample modal features of each modality and the labeled copyright information.
[0196] A predicted modal weight determination subunit, configured to determine the predicted modal weight of each modality based on the correlation coefficient.
[0197] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.
[0198] As can be seen from the above, in the embodiment of the present application, the acquisition unit 301 acquires the content to be detected and the source content set, and the source content set includes at least one source content with copyright; the feature extraction unit 302 performs multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and performs multi-modal feature extraction on the source content to obtain the source modal features of each modality; the calculation unit 303 calculates the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality; the determination unit 304 determines the modal weight of each modality according to the to-be-detected modal features, and weights the modal similarity based on the modal weight; the detection unit 305 detects the copyright information of the content to be detected in the source content set based on the weighted modal similarity. In this way, by determining the modal weight of each modality to weight the modal similarity of the content to be detected, the influence of the more important modality in different contents to be detected on the detection result is increased, and at the same time, the influence of the noise artificially introduced in the content to be detected on the accuracy of the detection result is reduced. Furthermore, the copyright information of the content to be detected is detected in the source content set according to the weighted modal similarity, so as to improve the accuracy of the content detection result and further improve the content detection efficiency.
[0199] The embodiment of the present application also provides a computer device, as Figure 6 shown, which shows a schematic structural diagram of the computer device involved in the embodiment of the present application. The computer device may be a server. Specifically:
[0200] The computer device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404 and other components. Those skilled in the art can understand that Figure 6 the computer device structure shown in
[0201] does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine some components, or arrange different components. Among them:
[0202] The memory 402 can be used to store software programs and modules. By running the software programs and modules stored in the memory 402, the processor 401 can execute various functional applications and content detection. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 402 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.
[0203] The computer device further includes a power supply 403 for powering each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0204] The computer device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0205] Although not shown, the computer device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to realize various functions as follows:
[0206] Obtain the content to be detected and a source content set, where the source content set includes at least one copyrighted source content; perform multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and perform multi-modal feature extraction on the source content to obtain the source modal features of each modality; calculate the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality; determine the modal weight of each modality according to the to-be-detected modal features, and weight the modal similarity based on the modal weight; based on the weighted modal similarity, detect the copyright information of the content to be detected in the source content set.
[0207] For the specific implementation of each of the above operations, reference may be made to the foregoing embodiments, which will not be elaborated herein. It should be noted that the computer device provided in the embodiments of the present application and the content detection method in the foregoing embodiments belong to the same concept. For the specific implementation process, refer to the foregoing method embodiments, which will not be elaborated herein.
[0208] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0209] Therefore, an embodiment of the present application provides a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any content detection method provided in the embodiments of the present application. For example, the instructions can execute the following steps:
[0210] Obtain the content to be detected and a set of source contents, where the set of source contents includes at least one source content with copyright; perform multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and perform multi-modal feature extraction on the source content to obtain the source modal features of each modality; calculate the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality; determine the modal weight of each modality based on the to-be-detected modal features, and weight the modal similarity based on the modal weight; detect the copyright information of the content to be detected in the set of source contents based on the weighted modal similarity.
[0211] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0212] Since the instructions stored in the computer-readable storage medium can execute the steps in any content detection method provided in the embodiments of the present application, the beneficial effects that can be achieved by any content detection method provided in the embodiments of the present application can be realized. For details, refer to the foregoing embodiments, which will not be elaborated herein.
[0213] Among them, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementation manners provided in the above embodiments.
[0214] The above has introduced in detail a content detection method, device, and computer-readable storage medium provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A content detection method, characterized in that, Including: Obtain the content to be detected and a set of source contents, where the set of source contents includes at least one source content with copyright; Perform multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and perform multi-modal feature extraction on the source content to obtain the source modal features of each modality; Calculate the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality; Obtain at least one content sample pair, where the content sample pair includes a detected content sample, a source content sample, and labeled copyright information; Use the weight determination sub-model in the preset content detection model to perform multi-modal feature extraction on the detected content sample to obtain the sample modal features of each modality; Calculate the correlation coefficient between the sample modal features of each modality and the labeled copyright information; Based on the correlation coefficient, determine the predicted modal weight of each modality; Based on the predicted modal weight, predict the modal similarity of the content sample pair to obtain the predicted modal similarity of each modality; According to the predicted modal similarity and the labeled copyright information, converge the preset content detection model to obtain a trained content detection model; According to the to-be-detected modal features, use the trained content detection model to determine the modal weight of each modality, and weight the modal similarity based on the modal weight; Based on the weighted modal similarity, detect the copyright information of the content to be detected in the set of source contents.
2. The content detection method according to claim 1, wherein Also including: Identify the modal importance score corresponding to the to-be-detected modal features; Detect the modal quality score of the to-be-detected modal features; Fuse the modal importance score and the modal quality score of the corresponding modality to obtain the modal weight of each modality.
3. The content detection method according to claim 1, wherein The detecting the copyright information of the content to be detected in the set of source contents based on the weighted modal similarity includes: Fuse the weighted modal similarities corresponding to each modality to obtain the total weighted modal similarity corresponding to each source content; According to the total weighted modal similarity, screen out the target source content in the set of source contents; Based on the target source content, determine the copyright information of the content to be detected.
4. The content detection method according to claim 1, characterized in that, The performing multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and performing multi-modal feature extraction on the source content to obtain the source modal features of each modality includes: Perform multi-modal extraction on the content to be detected to obtain the to-be-detected modal data of each modality, and perform multi-modal extraction on the source content to obtain the source modal data of each modality; According to the modal type corresponding to each modality, respectively determine the encoders corresponding to the to-be-detected modal data and the source modal data; Based on the encoder corresponding to the to-be-detected modal data, perform feature extraction on the to-be-detected modal data to obtain the to-be-detected modal features of each modality; Based on the encoder corresponding to the source modal data, perform feature extraction on the source modal data to obtain the source modal features of each modality.
5. The content detection method according to claim 1, wherein Converging the preset content detection model according to the predicted modal similarity and the labeled copyright information to obtain a trained content detection model includes: Calculating the total predicted modal similarity of the content sample pairs according to the predicted modal similarity of each modality; Determining the target loss information of each content sample pair based on the total predicted modal similarity and the corresponding labeled copyright information; Converging the preset content detection model based on the target loss information to obtain a trained content detection model.
6. A content detection device, characterized in that, Including: An acquisition unit configured to acquire the content to be detected and a source content set, where the source content set includes at least one source content with copyright; A feature extraction unit configured to perform multi-modal feature extraction on the content to be detected to obtain the to-be-detected modal features of each modality, and perform multi-modal feature extraction on the source content to obtain the source modal features of each modality; A calculation unit configured to calculate the similarity between the to-be-detected modal features and the source modal features of the corresponding modality to obtain the modal similarity of each modality; A determination unit configured to obtain at least one content sample pair, where the content sample pair includes a detected content sample, a source content sample, and labeled copyright information; Performing multi-modal feature extraction on the detected content sample by using a weight determination sub-model in the preset content detection model to obtain the sample modal features of each modality; Calculating the correlation coefficient between the sample modal features of each modality and the labeled copyright information; determining the predicted modal weight of each modality based on the correlation coefficient; predicting the modal similarity of the content sample pair based on the predicted modal weight to obtain the predicted modal similarity of each modality; converging the preset content detection model according to the predicted modal similarity and the labeled copyright information to obtain a trained content detection model; determining the modal weight of each modality by using the trained content detection model according to the to-be-detected modal features, and weighting the modal similarity based on the modal weight; A detection unit configured to detect the copyright information of the content to be detected in the source content set based on the weighted modal similarity.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the content detection method according to any one of claims 1 to 5.
8. A computer device, characterized in that, Including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor executes the computer program to implement the content detection method according to any one of claims 1 to 5.
9. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a storage medium, and a processor of a computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions to enable the computer device to execute the content detection method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Topic retrieval method and system based on multi-modal cross comparison
CN113392196A
Content retrieval method and device and computer readable storage medium
CN113821687A