A method and related device for deduplicating media content
By embedding subject features into feature vectors using computer vision and machine learning, the method improves media content duplication detection, ensuring accurate differentiation and increasing the quantity of usable content in recommendation systems.
Patent Information
- Application Number
- CN202110368996.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-04-06
AI Technical Summary
In the prior art, due to the large loss of feature vector information in media content, it is difficult to reflect the differences in media content with similar backgrounds but different subjects, resulting in frequent occurrence of misduplication and reducing content supply.
By identifying the media content subjects, embedding the subject features into the feature vector, enhancing the weight of the subject in the final image feature vector, and using computer vision and machine learning technology to extract and splice the subject feature vectors to improve the accuracy of differential recognition.
Effectively reduce the weight of mistakes, increase the content supply in the media content recommendation pool, improve the accuracy of content differences identification, and save resources.
Smart Images

Figure CN113704506B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and particularly to a method for deduplicating media content and related devices. Background Art
[0002] In the era of the rapid development of the Internet, with the lowering of the threshold for media content production, the upload and release volume of media content has grown exponentially. As content producers, users can upload media content on new media platforms to attract the attention of other users and bring huge traffic to the platforms. In particular, high-quality content producers, including the high-quality content behind them, have become the targets pursued by these platforms. As content producers, users can also obtain benefits through revenue sharing or rewards.
[0003] In order to increase their income, content creators will upload a large number of similar media content. Taking videos as an example of media content, content creators simply edit and modify the videos or directly copy and plagiarize the duplicate content of other account owners. As a result, the plagiarized content prevents the normal content of account owners from being enabled, and at the same time occupies a large amount of traffic, which is not conducive to the healthy development of the entire content ecosystem. Therefore, deduplication of media content is an important link.
[0004] In related technologies, mainly by directly extracting feature vectors from the images of different media content, and then determining whether different media content is similar according to the feature vectors, and then performing deduplication.
[0005] However, in this deduplication method, due to a large amount of information loss in the extracted feature vectors, it is difficult to better reflect the differences between different media content with similar backgrounds but different actual contents, and thus there is a situation of incorrect deduplication, reducing the content supply in the media content recommendation pool. Summary of the Invention
[0006] In order to solve the above technical problems, this application provides a method for deduplicating media content and related devices, which can effectively reduce the amount of incorrect deduplication, increase the enabled amount of content in the media content recommendation pool, and enrich the content supply of media content.
[0007] The embodiments of this application disclose the following technical solutions:
[0008] On the one hand, the embodiments of this application provide a method for deduplicating media content, and the method includes:
[0009] Obtain a first image set corresponding to the first media content and a second image set corresponding to the second media content;
[0010] Extract a first feature vector from the first image in the first image set, and extract a second feature vector from the second image in the second image set;
[0011] Perform subject recognition on the first image in the first image set to obtain first subject features, and perform subject recognition on the second image in the second image set to obtain second subject features;
[0012] Concatenate the first subject features and the first feature vector belonging to the same first image to obtain a first target feature vector corresponding to the first image, and concatenate the second subject features and the second feature vector belonging to the same second image to obtain a second target feature vector corresponding to the second image;
[0013] If it is determined that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector, perform duplicate removal processing.
[0014] On the other hand, an embodiment of the present application provides a media content duplicate removal device, which includes an acquisition unit, an extraction unit, an identification unit, a concatenation unit, and a duplicate removal unit:
[0015] The acquisition unit is configured to acquire a first image set corresponding to first media content and a second image set corresponding to second media content;
[0016] The extraction unit is configured to perform feature extraction on the first image in the first image set to obtain a first feature vector, and perform feature extraction on the second image in the second image set to obtain a second feature vector;
[0017] The identification unit is configured to perform subject recognition on the first image in the first image set to obtain first subject features, and perform subject recognition on the second image in the second image set to obtain second subject features;
[0018] The concatenation unit is configured to concatenate the first subject features and the first feature vector belonging to the same first image to obtain a first target feature vector corresponding to the first image, and concatenate the second subject features and the second feature vector belonging to the same second image to obtain a second target feature vector corresponding to the second image;
[0019] The duplicate removal unit is configured to perform duplicate removal processing if it is determined that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector.
[0020] On the other hand, an embodiment of the present application provides a device for media content duplicate removal, which includes a processor and a memory:
[0021] The memory is used to store program code and transmit the program code to the processor;
[0022] The processor is configured to execute the method described in the above aspects according to the instructions in the program code.
[0023] On the other hand, an embodiment of the present application provides a computer-readable storage medium for storing a computer program for executing the method described in the above aspects.
[0024] On the other hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described in the above aspects.
[0025] As can be seen from the above technical solutions, after the user uploads media content, in order to determine whether there is content duplication between the uploaded media contents, that is, whether there are behaviors such as copying and plagiarism. Taking the first media content and the second media content in the uploaded media contents as examples, the first image set corresponding to the first media content and the second image set corresponding to the second media content can be obtained. Feature extraction is performed on the first images in the first image set to obtain first feature vectors, and feature extraction is performed on the second images in the second image set to obtain second feature vectors. Since there may be a relatively large area of background similarity but differences in the main bodies between some media contents, in order to more accurately reflect the differences between these media contents, the first images in the first image set can be further subject-identified to obtain first subject features, and the second images in the second image set can be subject-identified to obtain second subject features. Then, the first subject feature and the first feature vector belonging to the same first image are spliced to obtain a first target feature vector corresponding to the first image, and the second subject feature and the second feature vector belonging to the same second image are spliced to obtain a second target feature vector corresponding to the second image, thereby embedding the subject feature into the feature vector of the image, which is equivalent to enhancing the weight of the subject in the final image feature vector, making the media contents of different subjects more different and more accurately reflecting the differences between these media contents. Whether the first media content and the second media content are similar is determined according to the first target feature vector and the second target feature vector obtained in this way, and then duplicate removal processing is performed when the two are similar, which can effectively reduce the amount of incorrect duplicate removal, increase the activation amount of the content in the media content recommendation pool, and enrich the supply amount of media content. Description of the Drawings
[0026] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 The implementation process of a media content deduplication method provided by the related art;
[0028] Figure 2 An example diagram of a media content provided by an embodiment of the present application;
[0029] Figure 3 A schematic diagram of an application scenario of a media content deduplication method provided by an embodiment of the present application;
[0030] Figure 4 A schematic diagram of the process of a media content deduplication method provided by an embodiment of the present application;
[0031] Figure 5 A schematic diagram of embedding the main body features into the feature vector using a feature matching model provided by an embodiment of the present application;
[0032] Figure 6 A schematic diagram of the process of a feature matching model training method provided by an embodiment of the present application;
[0033] Figure 7 Example diagrams of adjacent video frames of different videos provided by an embodiment of the present application;
[0034] Figure 8 A schematic diagram of the training process in which the training of the regression model and the training of the feature matching model are alternately performed provided by an embodiment of the present application;
[0035] Figure 9a An example diagram of an encoding method provided by an embodiment of the present application;
[0036] Figure 9b An example diagram of a tanh activation function provided by an embodiment of the present application;
[0037] Figure 10 An example diagram of a sign function provided by an embodiment of the present application;
[0038] Figure 11 A schematic diagram of the structure of a media content deduplication system provided by an embodiment of the present application;
[0039] Figure 12 A schematic diagram of the structure of a media content deduplication device provided by an embodiment of the present application;
[0040] Figure 13 A structural schematic diagram of the terminal device provided by an embodiment of the present application;
[0041] Figure 14 A structural schematic diagram of the server provided by an embodiment of the present application. Detailed implementation manners
[0042] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0043] In the related art, mainly by directly extracting feature vectors from images of different media contents, and then determining whether different media contents are similar according to the feature vectors, and then performing duplicate removal. See Figure 1 As shown, taking the media content including video 101 as an example, video frames are extracted for the video, and the image feature 102 corresponding to each video frame is obtained. Then, the image feature is compared with the video fingerprints of other videos in the video fingerprint library 103 to determine the similar video frames 104, and further determine whether the video is similar to other videos, so as to obtain all similar videos 105.
[0044] However, in some media contents, such as a large number of media contents corresponding to lectures, weather forecasts, news broadcasts, etc., different people are often in similar backgrounds. At this time, a corresponding video frame is, for example Figure 2 As shown. Figure 2 The picture on the left (when the media content includes a video, a video frame extracted), and Figure 2 The picture on the right (when the media content includes a video, a video frame extracted).
[0045] Due to the large area of similar backgrounds, although the people are different, the feature vectors extracted in the related art are difficult to reflect this difference. Furthermore, Figure 2 The picture on the left and Figure 2 The picture on the right are recognized as similar pictures, resulting in incorrect duplicate removal and reducing the content supply in the media content recommendation pool.
[0046] Therefore, the present application provides a media content duplicate removal method and related device. When extracting the feature vectors of media contents, the main body features in the media contents can be embedded into the original feature vectors to obtain the final feature vectors, that is, target feature vectors. The target feature vectors determined by this method are equivalent to increasing the weight of the main body in the final feature vectors, which will make the media contents of different main bodies more different and more accurately reflect the differences between these media contents. For media contents with similar large-area backgrounds but different main bodies, such as Figure 2 The picture on the left and Figure 2 The picture on the right, due to the addition of the main body features, since Figure 2 The picture on the left andFigure 2 If the people in the pictures on the right are different, then the target feature vectors determined by the method provided in the embodiments of the present application will clearly reflect this difference, avoiding determining the two as similar pictures due to the similarity of large areas of the background, thereby effectively reducing the amount of incorrect duplicate removal, increasing the activation amount of the content in the media content recommendation pool, and enriching the supply amount of media content.
[0047] The media content duplicate removal method provided in the embodiments of the present application is implemented based on artificial intelligence. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0048] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0049] In the embodiments of the present application, the artificial intelligence software technologies mainly involved include the above-mentioned computer vision, machine learning / deep learning, etc.
[0050] Computer Vision Technology (CV): Computer vision is a science that studies how to enable machines to "see". Further, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0051] Embodiments of the present application may, for example, be related to image recognition (IR), image semantic understanding (ISU), video processing, etc. in computer vision. Among them, image recognition can mainly be used for similar image detection / deduplication; image semantic understanding is mainly used for image feature extraction, including extraction of a first feature vector, a second feature vector, a first subject feature, and a second subject feature; video processing is mainly used for frame processing of videos and extraction of video frames, etc.
[0052] Machine learning is a multi-disciplinary cross-cutting subject, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks. Through machine learning and deep learning, various neural network models can be trained, so as to predict whether the first media content is similar to the second media content according to the neural network model, and further achieve deduplication of media content.
[0053] The media content deduplication method provided by the present application can be applied to media content deduplication devices with data processing capabilities, such as terminal devices and servers. Among them, the terminal device can specifically be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc.; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.
[0054] It should be noted that the method provided by the embodiments of the present application can be applied to new media platforms. After a user uploads media content to a new media platform, the media content uploaded by the user can be retrieved and matched to confirm whether there is content similarity between the media contents, so as to confirm whether the user has behaviors such as copying or plagiarism, such as simply editing and modifying the media content of oneself or other users, such as video titles, watermarks, or editing and cropping, adding advertisement opening and ending credits, modifying audio, etc., or directly copying and plagiarizing the duplicate content of other users, so as to perform deduplication of media content.
[0055] To facilitate the understanding of the technical solution of this application, the media content deduplication method provided in the embodiments of this application will be introduced below in combination with an actual application scenario, taking a server as the media content deduplication device.
[0056] See Figure 3 , Figure 3 which is a schematic diagram of the application scenario of the media content deduplication method provided in the embodiments of this application. In the Figure 3 shown application scenario, it includes a server 301 and a terminal device 302 used by a user. Among them, the server 301 serves as the aforementioned media content deduplication device.
[0057] In actual applications, the user can use a registered self-media account on the terminal device 302 to publish media content within a new media platform. The server 301 can obtain the media content published by the user through the network. The media content can be, for example, articles, videos, pictures, etc. that include pictures or videos.
[0058] To determine whether there is content duplication between the uploaded media content, taking the first media content and the second media content in the uploaded media content as an example, the server 301 can obtain the first image set corresponding to the first media content and the second image set corresponding to the second media content.
[0059] Among them, the first image set is an image set determined according to the videos or images included in the first media content, including at least one first image; the second image set is an image set determined according to the videos or images included in the second media content, including at least one second image. Usually, the second media content is the media content that has been uploaded by user A or uploaded simultaneously with the first media content or uploaded by other users. If the first media content includes a video, the second media content is the media content that includes a video; if the first media content includes a picture, the second media content is the media content that includes a picture.
[0060] The server 301 extracts features from the first images in the first image set to obtain first feature vectors, and extracts features from the second images in the second image set to obtain second feature vectors. Since there may be a relatively large area of background similarity but differences in the main subjects among some media contents, in order to more accurately reflect the differences between these media contents, the server 301 can further perform main subject recognition on the first images in the first image set to obtain first main subject features, and perform main subject recognition on the second images in the second image set to obtain second main subject features. Then, the first main subject features and the first feature vectors belonging to the same first image are spliced to obtain a first target feature vector corresponding to the first image, and the second main subject features and the second feature vectors belonging to the same second image are spliced to obtain a second target feature vector corresponding to the second image, thereby embedding the main subject features into the feature vectors of the images. This is equivalent to enhancing the weight of the main subject in the final image feature vectors, making the differences between media contents of different main subjects greater and more accurately reflecting the differences between these media contents.
[0061] The server 301 determines whether the first media content is similar to the second media content based on the obtained first target feature vector and second target feature vector. If they are similar, deduplication processing is performed. If they are not similar, the first media content and the second media content can be saved in the media content recommendation pool.
[0062] In Figure 3 taking the first media content and the second media content as pictures as an example, the first media content shown in 303 and the second media content shown in 304 have similar backgrounds but different characters. Compared with the related technologies, the method provided in the embodiments of the present application can distinguish two pictures with similar large-area backgrounds but different main subjects by embedding the main subject features into the feature vectors of the images, enhancing the weight of the main subject in the final image feature vectors, thereby determining that they are not similar. Furthermore, it allows the first media content to be put into the media recommendation pool, instead of determining that they are similar as in the related technologies, avoiding the amount of incorrect deduplication, increasing the activation amount of the content in the media content recommendation pool, and enriching the supply amount of media content.
[0063] The following specifically introduces the media content deduplication method provided in the embodiments of the present application with the server as the media content deduplication device.
[0064] See Figure 4 which Figure 4 is a flowchart of a media content deduplication method provided in the embodiments of the present application. As shown in Figure 4 the media content deduplication method includes the following steps:
[0065] S401. Obtain a first image set corresponding to the first media content and a second image set corresponding to the second media content.
[0066] In the new media era, a platform that enables users to voice their opinions, share, complain, and spread information by themselves is called "We Media". Users can use terminal programs and / or server-side programs to publish media content on the We Media platform through their We Media accounts. For the first media content and the second media content uploaded to the We Media platform, duplicate checking can be performed on the first media content and the second media content.
[0067] The first media content and the second media content refer to the media content published by users on the We Media platform through their We Media accounts. The media content published on the We Media platform can be provided for other users to view, and its display forms include but are not limited to articles, pictures, and videos. Among them, an article may include any one or a combination of pictures and videos, and a video includes a vertical video and a horizontal video. Users can publish it on the We Media platform through their We Media accounts and provide it for other users on the platform to view in the form of a Feeds stream.
[0068] It should be noted that Feeds, also translated as source material, feed, information provision, contribution, abstract, source, news subscription, web source (English: web feed, news feed, syndicated feed), is a data format through which a website spreads the latest information to users, usually arranged in a timeline (Timeline) manner. The timeline is the most primitive, intuitive, and basic display form of Feeds. The prerequisite for users to be able to subscribe to a website is that the website provides a source of information. Gathering Feeds in one place is called aggregation, and the software used for aggregation is called an aggregator. For end-users, an aggregator is software specifically used to subscribe to websites, and is generally also called an RSS reader (Rich Site Summary Reader), feed reader, news reader, etc.
[0069] It can be understood that the first image set is an image set determined according to the videos or images included in the first media content, including at least one first image; the second image set is an image set determined according to the videos or images included in the second media content, including at least one second image. If the first media content and the second media content include pictures, the first image in the first image set and the second image in the second image set are the pictures themselves. If the first media content and the second media content include videos, the first image in the first image set and the second image in the second image set are video frames extracted from the videos.
[0070] In the case where the first media content and the second media content include videos, multiple first video frames can be extracted from the first media content to obtain a first image set. The first media content is represented by the multiple first video frames, and the first images in the first image set are arranged in the time sequence of the multiple first video frames in the first media content. Multiple second video frames are extracted from the second media content to obtain a second image set. The second media content is represented by the multiple second video frames, and the second images in the second image set are arranged in the time sequence of the multiple second video frames in the second media content. Herein, the number of first video frames in the first image set may be the same as or different from the number of second video frames in the second image set, and this embodiment does not limit this.
[0071] It should be noted that the method of extracting multiple first video frames from the first media content and multiple second video frames from the second media content can be random extraction or extracting one frame every preset time interval (such as 0.1 s). Of course, considering the computational complexity and cost comprehensively, the number of extracted video frames can be limited not to exceed a preset threshold (including a first preset threshold and a second preset threshold), and the preset threshold is, for example, 30 frames. For a video with a long duration, such as a video exceeding 30 seconds, key frames of the video can be preferentially extracted. If the number of key frames is less than 30, frames are evenly extracted before and after the key frames to make up the deficiency.
[0072] In this case, a possible implementation method of extracting video frames is to extract first key video frames from the first media content. If the number of first key video frames is less than the first preset threshold, the video frames before and after the first key video frames in the first media content are evenly extracted until the total number of extracted video frames reaches the first preset threshold, thereby obtaining a first image set. Second key video frames are extracted from the second media content. If the number of second key video frames is less than the second preset threshold, the video frames before and after the second key video frames in the second media content are evenly extracted until the total number of extracted video frames reaches the second preset threshold, thereby obtaining a second image set. Herein, the magnitudes of the first preset threshold and the second preset threshold may be the same or different, and this embodiment does not limit this.
[0073] S402. Extract feature vectors from the first images in the first image set, and extract feature vectors from the second images in the second image set.
[0074] In order to compare whether the first media content and the second media content are similar and thus achieve deduplication of media content, feature vectors can be extracted from the first images in the first image set and feature vectors can be extracted from the second images in the second image set.
[0075] It should be noted that in this embodiment, the first feature vector and the second feature vector can be extracted through a feature matching model. For example, the first image in the first image set and the second image in the second image set are input into the feature matching model. Taking any first image or second image as an example, see Figure 5 As shown, specifically, a branch in the feature matching model, such as a feature extraction sub-model, can be used to determine the first feature vector or the second feature vector. Here, the feature extraction sub-model can be any neural network model that extracts image feature vectors. For example, it can be a Residual Network (ResNet), such as ResNet50, ResNet101, etc.
[0076] S403. Perform subject recognition on the first image in the first image set to obtain a first subject feature, and perform subject recognition on the second image in the second image set to obtain a second subject feature.
[0077] S404. Concatenate the first subject feature and the first feature vector belonging to the same first image to obtain a first target feature vector corresponding to the first image, and concatenate the second subject feature and the second feature vector belonging to the same second image to obtain a second target feature vector corresponding to the second image.
[0078] For media content with similar large-area backgrounds but different subjects, the subject features can better reflect the differences between media contents. Therefore, subject object detection can be introduced to extract subject features, and then the subject features can be embedded into the previously extracted feature vectors to jointly represent the media content.
[0079] Based on this, in this embodiment, subject recognition can be performed on the first image in the first image set to obtain a first subject feature, and subject recognition can be performed on the second image in the second image set to obtain a second subject feature. The first subject feature and the first feature vector belonging to the same first image are concatenated to obtain a first target feature vector corresponding to the first image, and the second subject feature and the second feature vector belonging to the same second image are concatenated to obtain a second target feature vector corresponding to the second image.
[0080] It should be noted that in this embodiment, the first subject feature and the second subject feature can be extracted through a feature matching model. For example, the first image in the first image set and the second image in the second image set are input into the feature matching model, and the subject detection sub-model in the feature matching model is used to extract the first subject feature and the second subject feature. Then, the first subject feature and the first feature vector belonging to the same first image are concatenated through the concatenation layer in the feature matching model to obtain a first target feature vector, and the second subject feature and the second feature vector belonging to the same second image are concatenated through the concatenation layer in the feature matching model to obtain a second target feature vector. This is equivalent to enhancing the weight of the subject in the embedding of the final image, which will make the embeddings of images including different subjects more different.
[0081] Taking any one of the first images or the second images as an example, refer to Figure 5 as shown. Specifically, another branch in the feature matching model, such as the subject detection sub-model, can be used to determine the first subject feature or the second subject feature. Then, the subject feature is embedded into the feature vector. Taking Figure 5 the image shown as any one of the first images as an example, after obtaining the first feature vector and the first subject feature, the first feature vector and the first subject feature can be concatenated through the concatenation layer to obtain a first target feature vector. The finally used first target feature vector can be obtained after encoding, such as using hash encoding, so as to be used as the fingerprint of this first image.
[0082] It can be understood that the subject detection sub-model can be various models for detecting subject targets, such as the YOLO model or the Single Shot MultiBox Detector (SSD) model, etc. In this embodiment, the YOLO model is used to detect the coordinates of the subject in the image (including the first image and the second image) and determine the subject feature. Among them, the detected subject can be all or part of the subjects in the image, and the part of the subject can be the largest subject, etc. This embodiment does not make any limitations in this regard.
[0083] The YOLO (You Only Look Once) model is an object recognition and localization algorithm based on deep neural networks. Its biggest feature is its very fast running speed, which can be used in real-time systems, has high efficiency in detecting the main bodies of a large number of images, and has a lower overall machine cost. Now the YOLO model has evolved to version v5, but the new version is also continuously improved and evolved based on the original version. SSD is an object detection algorithm and is currently one of the main detection frameworks, mainly used to solve the problem of main body detection (localization + classification), that is, input an image and output the position information and category information of multiple boxes. Differences between SSD / YOLO: YOLO connects a fully connected layer after the convolutional layer, that is, only the highest-level feature maps (Featuremaps) are used during detection. SSD adopts a pyramid structure, that is, it uses these feature maps of different sizes such as conv4-3 / conv-7 / conv6-2 / conv7-2 / conv8_2 / conv9_2, and performs softmax classification and position regression simultaneously on multiple feature maps. SSD also adds Prior box. The PriorBox layer in the SSD network is used to deploy the default boxes at each position (pixel point) in the feature map.
[0084] S405. If it is determined that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector, perform duplicate removal processing.
[0085] After obtaining the first target feature vector that can accurately represent the first media content and the second target feature vector that can accurately represent the second media content, determine whether the first media content is similar to the second media content according to the first target feature vector and the second target feature vector, and thus perform duplicate removal processing according to the similarity determination result.
[0086] It should be noted that in this embodiment, only the first media content and the second media content are taken as examples. In fact, all similar media content can be retrieved through the above method, and then duplicate removal is performed on the similar media content.
[0087] In this embodiment, if the feature matching model includes a matching sub-model, it is possible to determine whether the first media content is similar to the second media content according to the first target feature vector and the second target feature vector through the matching sub-model in the feature matching model.
[0088] It should be noted that in this embodiment, it is possible to determine whether the first media content is similar to the second media content through the Faiss vector retrieval method. Faiss is an open-source clustering and similarity search library developed by the Facebook AI team. It provides efficient similarity search and clustering for dense vectors, supports searches for vectors on the order of billions, and is currently the most mature approximate nearest neighbor search library. It includes various algorithms for searching vector sets of any size, as well as support code for algorithm evaluation and parameter adjustment. The Faiss library contains multiple methods for similarity search, and its core modules include high-performance clustering, principal component analysis (PCA), and product quantization (PQ). It assumes that instances are represented as vectors and identified by integers, and at the same time, vectors can be compared with feature distances or dot products. Vectors similar to the query vector are those with the lowest feature distance (such as L2 distance) or the highest dot product with the query vector. It also supports cosine similarity, that is, the similarity calculation methods adopted in Faiss are mainly two: Euclidean distance and dot product. In this embodiment of the present application, the Euclidean distance is mainly used as an example for introduction. Once these vectors are extracted by the learning machine (from images, videos, text files, or other channels), they can already be input into the similarity search library for retrieval and matching. In this embodiment, the first target feature vector corresponding to the first media content is extracted and input into the similarity search library for retrieval and matching. During the retrieval and matching process, it is also necessary to determine the second target feature vector corresponding to the second media content (media content in the similarity search library), so as to determine the similarity between the first media content and the second media content according to the first target feature vector and the second target feature vector. If the similarity meets the preset conditions, it is determined that the first media content is similar to the second media content.
[0089] It can be understood that in this embodiment of the present application, the first media content and the second media content may include pictures or videos. When the first media content and the second media content include different types of content, the method for calculating the similarity may also be different.
[0090] If the first media content and the second media content include pictures, the way to determine the similarity between the first media content and the second media content can be to determine the second feature distance between the first media content and the second media content according to the first target feature vector and the second target feature vector. The second feature distance is used to represent the similarity between the first media content and the second media content. In this case, if the similarity meets the preset conditions, the way to determine that the first media content is similar to the second media content can be that if the second feature distance is less than or equal to the first distance threshold, it is determined that the first media content is similar to the second media content. At this time, the preset condition is that the second feature distance is less than or equal to the first distance threshold.
[0091] Among them, the first distance threshold can be preset according to actual requirements. For example, usually, the first distance threshold can be set to 150.
[0092] If the first media content and the second media content include videos, the way to determine the similarity between the first media content and the second media content can be to align the first image in the first image set with the second image in the second image set to establish the corresponding relationship between the first image and the second image. Then, for each pair of the first image and the second image with the corresponding relationship, according to the first target feature vector and the second target feature vector, determine the third feature distance between the first image and the second image. If the third feature distance is less than or equal to the second distance threshold, determine that the first image and the second image are similar, and obtain the number of pairs of similar first images and second images. The number of pairs of similar first images and second images is used to represent the similarity between the first media content and the second media content. The second distance threshold can be preset according to actual requirements. For example, usually, the second distance threshold can be set to 150.
[0093] For example, the number of the first images in the first image set is 30 frames, which are arranged in sequence as a1, a2, a3, ……, a30 according to the time sequence, and the number of the second images in the second image set is 25 frames, which are arranged in sequence as b1, b2, b3, ……, b25 according to the time sequence. Then, when aligning the first image in the first image set with the second image in the second image set to establish the corresponding relationship between the first image and the second image, that is, a1 has a corresponding relationship with b1, a2 has a corresponding relationship with b2, a3 has a corresponding relationship with b3, ……, a25 has a corresponding relationship with b25, so as to compare whether the first image and the second image with the corresponding relationship are similar, and further determine the similarity between the first media content and the second media content.
[0094] Using the number of pairs of similar first images and second images to represent the similarity between the first media content and the second media content can include various ways. Specifically, one way can be to directly use the number of pairs of similar first images and second images as the similarity. In this way, if the number of pairs of similar first images and second images reaches the preset quantity, determine that the first media content and the second media content are similar. Another way can be to obtain the total number of pairs of the first image and the second image, and use the ratio of the number of pairs of similar first images and second images to the total number of pairs as the similarity between the first media content and the second media content. In this way, if the ratio reaches the preset ratio, determine that the first media content and the second media content are similar. The preset ratio can be preset according to actual needs. For example, it can be 80%.
[0095] Among them, the total number of pairs of the first image and the second image can be the minimum of the number of the first images and the number of the second images. For example, after aligning the first image in the first image set with the second image in the second image set as described above, the obtained corresponding relationships are that a1 has a corresponding relationship with b1, a2 has a corresponding relationship with b2, a3 has a corresponding relationship with b3, ……, a25 has a corresponding relationship with b25. Then the total number of pairs of the first image and the second image is 25, and 25 is the number of the second images.
[0096] It can be seen from the above technical solution that when a user uploads media content, in order to determine whether there is content duplication among the uploaded media content, that is, whether there are behaviors such as copying and plagiarism, taking the first media content and the second media content in the uploaded media content as examples, the first image set corresponding to the first media content and the second image set corresponding to the second media content can be obtained. Feature extraction is performed on the first images in the first image set to obtain first feature vectors, and feature extraction is performed on the second images in the second image set to obtain second feature vectors. Since there may be a relatively large area of background similarity but differences in the main bodies among some media content, in order to more accurately reflect the differences among these media content, the main body of the first images in the first image set can be recognized to obtain first main body features, and the main body of the second images in the second image set can be recognized to obtain second main body features. Then, the first main body feature and the first feature vector belonging to the same first image are spliced to obtain the first target feature vector corresponding to the first image, and the second main body feature and the second feature vector belonging to the same second image are spliced to obtain the second target feature vector corresponding to the second image. Thus, the main body feature is embedded into the feature vector of the image, which is equivalent to enhancing the weight of the main body in the final image feature vector, making the media content of different main bodies more different and more accurately reflecting the differences among these media content. Whether the first media content and the second media content are similar is determined according to the obtained first target feature vector and second target feature vector, and then duplicate removal processing is performed when the two are similar, which can effectively reduce the amount of incorrect duplicate removal, increase the enabled amount of content in the media content recommendation pool, and enrich the supply of media content.
[0097] In addition, by embedding the main body feature into the feature vector, the effect of duplicate removal for pictures and videos is significantly improved, which can effectively reduce the unnecessary manual review and processing quantity during the information flow distribution process and save a large amount of resources.
[0098] Next, the training method of the feature matching model used in the above method will be introduced. Refer to Figure 6 The method includes:
[0099] S601. Obtain the third image set corresponding to the first historical media content in the training samples, and the fourth image set corresponding to the second historical media content in the training samples.
[0100] Among them, the first historical media content and the second historical media content are used as training samples for training the feature matching model, and it is known whether they are similar. Whether the first historical media content and the second historical media content are similar can be identified by the target label.
[0101] S602. Determine the third feature vectors corresponding to the images in the third image set and the fourth feature vectors corresponding to the images in the fourth image set through the feature extraction sub-model in the feature matching model.
[0102] S603. Determine the third subject features corresponding to the images in the third image set and the fourth subject features corresponding to the images in the fourth image set through the subject detection sub-model in the feature matching model.
[0103] S604. Concatenate the third subject features and the third feature vectors belonging to the same image through the concatenation layer in the feature matching model to obtain the third target feature vector, and concatenate the fourth subject features and the fourth feature vectors belonging to the same image through the concatenation layer in the feature matching model to obtain the fourth target feature vector.
[0104] In this embodiment, the subject features are embedded into the feature vectors during the training process. The processes of S602 - S604 are similar to the process of using the feature matching model for deduplication, and will not be elaborated here.
[0105] S605. Train the feature matching model according to the third target feature vector, the fourth target feature vector, and the target label.
[0106] The feature matching model predicts whether the first historical media content and the second historical media content are similar according to the third target feature vector and the fourth target feature vector, and adjusts the parameters of the feature matching model according to the prediction result and the target label to complete the training of the feature matching model.
[0107] In some cases, if the first historical media content and the second historical media content are videos, at this time, the images in the third image set are multiple video frames extracted from the first historical media content, and the images in the third image set are arranged in the time sequence of the video frames in the first historical media content; the images in the fourth image set are multiple video frames extracted from the second historical media content, and the images in the fourth image set are arranged in the time sequence of the video frames in the second historical media content. That is to say, in this embodiment, a video is represented by extracting multiple video frames. If multiple video frames are all similar, they will not be able to represent the entire video.
[0108] As Figure 7 shown, taking the third and fourth frames in Video A and Video B as examples, Figure 7 the lower left image in [[ID=]] should originally match and be similar to the lower right image, that is, the video frames corresponding to different videos are kept similar. However, Figure 7 the lower left image in [[ID=]] also matches and is similar to the upper right image, that is, multiple video frames extracted from the same video are similar, and thus it cannot represent the entire video, resulting in the determination result of whether the two videos are similar may be inaccurate.
[0109] To this end, in order to avoid the similarity of adjacent video frames of the same video caused by the introduction of the subject feature, which in turn affects the determination result of whether the two videos are similar, a strategy of maintaining the distance between adjacent video frames is introduced during the training process. During the training process, two input channels (pipelines) are set. One is the normal pipeline for contrast learning of positive and negative sample pairs, and the other is the pipeline prepared additionally for maintaining the distance between adjacent video frames taken from the same video. The distance maintenance strategy can be implemented through a regression model. The first feature distance between adjacent video frames in the third image set or the fourth image set is determined through the regression model, and the regression model is trained according to the first feature distance and the reference distance. The training of the regression model and the training of the feature matching model are carried out alternately. Among them, the reference distance can be the judgment of the distance between adjacent video frames by the model before the introduction of the subject feature.
[0110] The specific training process is as Figure 8 shown. The original data is read through the backbone network of the feature matching model for contrast learning. Each iteration can predict whether the images in the third image set are similar to the images in the fourth image set, and the prediction result is obtained. Among them, similarity can be represented by 1, and dissimilarity can be represented by 0. The cycle of alternating the training of the regression model and the training of the feature matching model can be represented by T, which means that in every T iterations, one iteration is to use the data of adjacent video frames for regression, and the remaining T - 1 iterations are all for contrast learning. In [[ID=]], it is equivalent to T = 4. Figure 8 In this embodiment, the loss function for training the regression model can be the L1 norm loss function (loss), which is to regress the feature distance of adjacent video frames to keep the two at a certain distance. The L1 norm loss function is also known as the Least Absolute Deviations (LAD) and the least absolute error (LAE). Generally speaking, it minimizes the sum of the absolute differences between the target value
[0111] and the estimated value where S is expressed as follows: the sum of the absolute differences is minimized, where, S is as follows:
[0112]
[0113] The binary classification similarity branch (HashNet Binary Loss) is the commonly used logarithmic loss function. For binary classification logistic regression, we have:
[0114]
[0115] θ represents the parameter vector, x is the input, T is the period, and hθ(x) is the binary classification branch, representing the output result to predict similarity. Its input range is , and for exactly (0, 1), it exactly meets the requirement of the probability distribution being (0, 1). It is convenient to describe the classifier with probability, which is much more convenient than simply using a certain threshold; it is a monotonically increasing function with good continuity and no discontinuity points.
[0116] Logarithmic loss function: , where X is the input, Y represents the prediction result obtained when the input is X, represents the probability of the prediction result Y obtained when the input is X.
[0117] In this embodiment, the logarithmic loss function is adopted. According to the foregoing content, the logarithmic likelihood loss function (cost function) of logistic regression can be obtained:
[0118]
[0119] Combining the above two expressions into one, the loss function of a single sample can be described as: , which is the final loss function expression of logistic regression, where y i is the input target prediction result.
[0120] It should be noted that the current backbone algorithm is a siamese network based on resnet50. The input image is used to extract feature vectors through the same siamese network and encoded into 01 vectors for dimensionality reduction. Specifically, as Figure 9a shown, the main purpose is to reduce the dimension, reduce the storage space, and at the same time not lose the progress, which is beneficial to large-scale engineering implementation. Figure 9a The encoding method shown is: scale transformation + tanh activation function + sign function.
[0121] The output of the tanh function is already relatively close to 01 (see Figure 9b), the output of tanh is subjected to loss constraint. Tanh is one of the hyperbolic functions, and tanh() is the hyperbolic tangent. In mathematics, the hyperbolic tangent "tanh" is derived from the basic hyperbolic functions hyperbolic sine and hyperbolic cosine, and the derivation formula is as follows:
[0122]
[0123] Then, the output of tanh is signed, and not much precision is lost. Sign is also called sgn, which means sign. The sign function (usually denoted as sign(x)) is a very useful type of function that can help us achieve some constructions that are difficult to directly implement in the Geometer's Sketchpad. The sign function can separate the sign of the function. In mathematics and computer operations, its function is to take the sign (positive or negative) of a certain number: when x > 0, sign(x) = 1; when x = 0, sign(x) = 0; when x < 0, sign(x) = -1, see Figure 10 as shown. In communication, sign(t) represents such a signal: when t ≥ 0, sign(t) = 1; that is, starting from the moment t = 0, the amplitude of the signal is 1; when t < 0, sign(t) = -1; before the moment t = 0, the amplitude of the signal is -1.
[0124] Here, the distance between two feature vectors is measured using the inner product of the feature vectors, and the range of the result is -1024 to 1024 (1024 dimensions). The Loss constraint is the contrastive loss function. When y = 1, the two are similar, and d is minimized as much as possible. When y = 0, Max(0, m - d) is minimized. m is the distance that dissimilar vectors should maintain. When d >= m, the loss is minimized. For example, if the distance between the 01 vectors here is measured using the Hamming distance and the m parameter is set to 15. The formula description is as follows:
[0125] Model: Convert the input data into a set of feature vectors. Here is the target feature vector obtained after splicing. The target feature vector includes the third target feature vector and the fourth target feature vector , or includes the first target feature vector and the first target feature vector .
[0126] Distance: The inner product between two 01 vectors.
[0127] Loss function: .
[0128] To better understand the media content deduplication method provided in the embodiments of the present application, the embodiments of the present application also provide a media content deduplication system. The media content deduplication system provided in the embodiments of the present application will be introduced below.
[0129] See Figure 11 , Figure 11 which is a schematic structural diagram of a media content deduplication system provided in the embodiments of the present application. As Figure 11 shown, the media content deduplication system includes a content production end 1101, a content consumption end 1102, an uplink and downlink content interface server 1103, a content database 1104, a scheduling center 1105, a manual review system 1106, a content storage server 1107, a download file system 1108, a frame extraction service 1109, a main body embedding vector generation service 1110, a distributed vector retrieval service 1111, a deduplication relationship chain calculation service 1112, and a content export distribution service 1113:
[0130] The content production end 1101 is used for:
[0131] (1) Content producers of Professional Generated Content (PGC), User Generated Content (UGC), Professional User Generated Content (PUGC), or Multi-Channel Network (MCN) provide media content through an Application Programming Interface (API), including graphic content or video content, which are the main content sources for distributed content;
[0132] (2) Upload media content through communication with the uplink and downlink content interface server 1103. If the media content includes graphic content, the source of the graphic content is usually a lightweight publishing end and an editing content entry; if the media content includes video content, the video content is usually published by a shooting and photography end, and music, filter templates, and video beautification functions that can be selected for the local video content during the shooting process, etc.
[0133] The content consumption end 1102 is used for:
[0134] (1) As a consumer, communicate with the uplink and downlink content interface server 1103, obtain the index information for accessing content through recommendations, and then communicate with the content storage server 1107 to obtain the corresponding content, including the content obtained through recommendations and the content subscribed to in special topics. The content storage server 1107 stores content entities such as video source files and picture source files, while the meta-information of the content, such as title, author, cover image, classification, tag information, etc., is stored in the content database 1104;
[0135] (2) At the same time, report the user's playback behavior data, such as stuttering, loading time, and playback during the upload and download processes, to the backend for statistical analysis;
[0136] (3) Usually browse content data through the Feeds stream method.
[0137] The uplink and downlink content interface server 1103
[0138] (1) The uplink and downlink content interface server 1103 communicates directly with the content production end 1101, and stores the media content submitted from the front end, usually the title, publisher, abstract, cover image, and release time of the media content, into the content database 1104;
[0139] (2) Write the meta-information of the media content, such as file size, cover image link, title, release time, author, resolution, bit rate, etc., into the content database 1104;
[0140] (3) Synchronize the submitted media content to the scheduling center 1105 for subsequent content processing and circulation.
[0141] The content database 1104 is used for:
[0142] (1) It is the core database of the content. All the meta-information of the content published by producers is stored in this business database. The key is the meta-information of the content itself, such as file size, cover image link, bit rate, file format, title, release time, author, the mark of whether it is original or the first release, and also includes the classification of the media content during the manual review process (including first, second, and third-level classifications and tag information. For example, for an article explaining mobile phone A, the first-level classification is technology, the second-level classification is smart phones, the third-level classification is domestic mobile phones, and the tag information is A, mate30);
[0143] (2) During the manual review process, the information in the content database 1104 will be read, and at the same time, the results and status of the manual review will also be transmitted back into the content database 1104;
[0144] (3) The content processing by the scheduling center 1105 mainly includes machine processing and manual review processing. The core of machine processing involves various quality judgments such as low-quality filtering, content tagging such as classification and tagging information, and content deduplication. Their results will be written into the content database 1104. Exactly the same duplicate content will not be processed manually again, which can effectively reduce the cost of manual processing.
[0145] (4) When extracting tags subsequently, the meta-information of the content will be read from the content database 1104.
[0146] The said scheduling center 1105 is used for:
[0147] (1) Responsible for the entire scheduling process of content transfer. It receives the incoming media content through the uplink and downlink content interface server 1103, and then obtains the meta-information of the media content from the content database 1104.
[0148] (2) Schedule the manual review system 1106 and the machine processing system, and control the order and priority of scheduling.
[0149] The said manual review system 1106 is used for:
[0150] (1) The content is enabled through the manual review system 1106, and then directly provided to the content consumption end 1102 through the content export distribution service 1113 (usually a recommendation engine or a search engine or operation). That is, the content index information obtained by the content consumption end 1102 is usually the entry uniform resource locator (URL) address for content access.
[0151] (2) The manual review system 1106 is the carrier of manual service capabilities, mainly used for reviewing and filtering content that is sensitive, pornographic, and not allowed by law and cannot be determined by machines. At the same time, it also performs label annotation on media content.
[0152] The said content storage server 1107 is used for:
[0153] (1) Store the content entity information other than the meta-information of the media content, such as the video source file and the picture source file of the graphic content.
[0154] (2) When extracting media content tags, provide the video source file including the frame extraction content in the source file for the tag service.
[0155] The said download file system 1108 is used for:
[0156] (1) Download and obtain the original media content from the content storage server 1107, and control the download speed and progress. Usually, it is a group of parallel servers, which are composed of relevant task scheduling and distribution clusters;
[0157] (2) Call the frame extraction service 1109 for the downloaded file to obtain the necessary image sets (such as the first image set and the second image set) from the source file, which serve as the data source for constructing feature vectors subsequently.
[0158] The frame extraction service 1109 is used for:
[0159] (1) Perform primary processing on the file downloaded from the content storage server 1107 by the download file system 1108 according to the algorithms and strategies mentioned above;
[0160] (2) In the case where the media content includes video, considering the calculation amount and cost comprehensively, at most 30 frames are extracted. For video content exceeding 30 seconds, key frames of the video are preferentially extracted. If there are less than 30 frames, frames are evenly extracted before and after the key frames to make up the number.
[0161] The main body embedding vector generation service 1110 is used for:
[0162] (1) According to the detailed algorithm model described above, construct the method for generating the main body embedding vector, train to obtain the corresponding feature matching model, and then construct the target feature vector embedding the main body features through this feature matching model;
[0163] (2) Provide the data source of the main body embedding vector together with the distributed vector retrieval service 1111.
[0164] The distributed vector retrieval service 1111 is used for:
[0165] (1) As described above, on the basis of the constructed main body embedding vector, perform distributed management and retrieval matching on the indexes of the vectors. Here, Faiss is specifically used to manage all the main body embedding vectors.
[0166] The duplicate relationship chain calculation service 1112 is used for:
[0167] (1) As described in detail above, after obtaining the target feature vector of the main body feature embedding, then through the distributed vector retrieval service 1111 and Figure 4 the duplicate elimination method provided in the corresponding embodiment, retrieve whether the first media content is repeated with the second media content;
[0168] (2) All the duplicate media content that meets the conditions retrieved is the result of the duplicate elimination calculation. At this time, it can be enabled according to product strategies such as the original account owner or the one with the highest quality and clarity.
[0169] To better understand the media content deduplication method provided in the embodiments of the present application, the above abnormal account determination process will be introduced below in combination with specific application scenarios.
[0170] The self-media platform calls the media content deduplication system to deduplicate the media content uploaded by users. Taking the media content as a video as an example, for a video (i.e., the first media content) and other videos (i.e., the second media content), frame extraction is performed to obtain a first image set and a first image set; then, the main body embedding feature vectors are generated through the main body embedding vector generation service, obtaining the first target feature vector corresponding to each video frame in the first image set and the second target feature vector corresponding to each video frame in the second image set. Furthermore, through the distributed vector retrieval service, it is determined whether the two videos are similar based on the first target feature vector and the second target feature vector, thereby retrieving all duplicate videos. Finally, according to product strategies, such as the original account owner or the one with the highest quality and clarity, it can be enabled.
[0171] For the media content deduplication method provided in the above embodiments, the embodiments of the present application also provide a media content deduplication device. Refer to Figure 12 , Figure 12 which is a structural diagram of a media content deduplication device provided in the embodiments of the present application. The device 1200 includes an acquisition unit 1201, an extraction unit 1202, an identification unit 1203, a splicing unit 1204, and a deduplication unit 1205:
[0172] The acquisition unit 1201 is configured to acquire a first image set corresponding to the first media content and a second image set corresponding to the second media content;
[0173] The extraction unit 1202 is configured to perform feature extraction on the first images in the first image set to obtain first feature vectors, and perform feature extraction on the second images in the second image set to obtain second feature vectors;
[0174] The identification unit 1203 is configured to perform main body identification on the first images in the first image set to obtain first main body features, and perform main body identification on the second images in the second image set to obtain second main body features;
[0175] The splicing unit 1204 is configured to splice the first main body features and the first feature vectors belonging to the same first image to obtain a first target feature vector corresponding to the first image, and splice the second main body features and the second feature vectors belonging to the same second image to obtain a second target feature vector corresponding to the second image;
[0176] The deduplication unit 1205 is configured to perform deduplication processing if it is determined that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector.
[0177] In a possible implementation, the first media content and the second media content include pictures, and the first picture in the first picture set and the second picture in the second picture set are the pictures themselves.
[0178] In a possible implementation, the first media content and the second media content include videos, and the obtaining unit 1201 is configured to:
[0179] Extract a plurality of first video frames from the first media content to obtain the first picture set, and the first pictures in the first picture set are arranged in the time sequence of the plurality of first video frames in the first media content;
[0180] Extract a plurality of second video frames from the second media content to obtain the second picture set, and the second pictures in the second picture set are arranged in the time sequence of the plurality of second video frames in the second media content; the number of first video frames in the first picture set is the same as the number of second video frames in the second picture set.
[0181] In a possible implementation, the obtaining unit 1201 is configured to:
[0182] Extract first key video frames from the first media content;
[0183] If the number of the first key video frames is less than a first preset threshold, uniformly extract the video frames before and after the first key video frames in the first media content until the total number of the extracted video frames reaches the first preset threshold to obtain the first picture set;
[0184] Extract second key video frames from the second media content;
[0185] If the number of the second key video frames is less than a second preset threshold, uniformly extract the video frames before and after the second key video frames in the second media content until the total number of the extracted video frames reaches the second preset threshold to obtain the second picture set.
[0186] In a possible implementation, the extraction unit 1202 is configured to determine the first feature vector and the second feature vector through a feature extraction sub-model in a feature matching model.
[0187] The recognition unit 1203 is configured to determine the first body feature and the second body feature through the body detection sub-model in the feature matching model;
[0188] The splicing unit 1204 is configured to splice the first body feature and the first feature vector belonging to the same first image through the splicing layer in the feature matching model to obtain the first target feature vector, and splice the second body feature and the second feature vector belonging to the same second image through the splicing layer in the feature matching model to obtain the second target feature vector;
[0189] The duplicate removal unit 1205 is configured to determine that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector through the matching sub-model in the feature matching model.
[0190] In a possible implementation manner, the apparatus further includes a training unit:
[0191] The training unit is configured to obtain a third image set corresponding to the first historical media content in the training sample, and a fourth image set corresponding to the second historical media content in the training sample, and whether the first historical media content and the second historical media content are similar is identified by a target label;
[0192] Determine a third feature vector corresponding to the image in the third image set and a fourth feature vector corresponding to the image in the fourth image set through the feature extraction sub-model in the feature matching model;
[0193] Determine a third body feature corresponding to the image in the third image set and a fourth body feature corresponding to the image in the fourth image set through the body detection sub-model in the feature matching model;
[0194] Splice the third body feature and the third feature vector belonging to the same image through the splicing layer in the feature matching model to obtain the third target feature vector, and splice the fourth body feature and the fourth feature vector belonging to the same image through the splicing layer in the feature matching model to obtain the fourth target feature vector;
[0195] Train the feature matching model according to the third target feature vector, the fourth target feature vector and the target label.
[0196] In a possible implementation, the first historical media content and the second historical media content are videos. The images in the third image set are multiple video frames extracted from the first historical media content, and the images in the third image set are arranged in the time sequence of the video frames in the first historical media content. The images in the fourth image set are multiple video frames extracted from the second historical media content, and the images in the fourth image set are arranged in the time sequence of the video frames in the second historical media content. The training unit is further configured to:
[0197] Determine a first feature distance between adjacent video frames in the third image set or the fourth image set through a regression model;
[0198] Train the regression model according to the first feature distance and a reference distance, and the training of the regression model and the training of the feature matching model are performed alternately.
[0199] In a possible implementation, the deduplication unit 1205 is configured to determine a similarity between the first media content and the second media content according to the first target feature vector and the second target feature vector;
[0200] If the similarity meets a preset condition, determine that the first media content and the second media content are similar.
[0201] In a possible implementation, if the first media content and the second media content include pictures, the deduplication unit 1205 is configured to determine a second feature distance between the first media content and the second media content according to the first target feature vector and the second target feature vector, and the second feature distance is used to represent the similarity between the first media content and the second media content;
[0202] If the second feature distance is less than or equal to a first distance threshold, determine that the first media content and the second media content are similar, and the preset condition is that the second feature distance is less than or equal to the first distance threshold.
[0203] In a possible implementation, if the first media content and the second media content include videos, the deduplication unit 1205 is configured to align a first image in the first image set with a second image in the second image set to establish a correspondence between the first image and the second image;
[0204] For each pair of the first image and the second image with a correspondence, determine a third feature distance between the first image and the second image according to the first target feature vector and the second target feature vector;
[0205] If the third feature distance is less than or equal to the second distance threshold, determine that the first image and the second image are similar;
[0206] Obtain the logarithm of the number of pairs of similar first images and second images, and the logarithm of the number of pairs of similar first images and second images is used to represent the similarity between the first media content and the second media content.
[0207] It can be seen from the above technical solution that when a user uploads media content, in order to determine whether there is content duplication between the uploaded media contents, that is, whether there are behaviors such as copying and plagiarism, taking the first media content and the second media content in the uploaded media contents as examples, the first image set corresponding to the first media content and the second image set corresponding to the second media content can be obtained. Feature extraction is performed on the first images in the first image set to obtain first feature vectors, and feature extraction is performed on the second images in the second image set to obtain second feature vectors. Since there may be a large area of background similarity but differences in the main bodies between some media contents, in order to more accurately reflect the differences between these media contents, the first images in the first image set can be further subjected to main body recognition to obtain first main body features, and the second images in the second image set can be subjected to main body recognition to obtain second main body features. Then, the first main body feature and the first feature vector belonging to the same first image are spliced to obtain the first target feature vector corresponding to the first image, and the second main body feature and the second feature vector belonging to the same second image are spliced to obtain the second target feature vector corresponding to the second image, so as to embed the main body feature into the feature vector of the image, which is equivalent to enhancing the weight of the main body in the final image feature vector, making the media contents of different main bodies more different and more accurately reflecting the differences between these media contents. Determine whether the first media content and the second media content are similar according to the first target feature vector and the second target feature vector obtained in this way, and then perform duplicate removal processing when the two are similar, which can effectively reduce the amount of incorrect duplicate removal, increase the activation amount of the content in the media content recommendation pool, and enrich the supply amount of media content.
[0208] The embodiment of the present application also provides a device for deduplicating media content. This device can be a terminal device. Taking the terminal device as a smart phone as an example:
[0209] Figure 13 The block diagram of a part of the structure of the smart phone related to the terminal device provided by the embodiment of the present application is shown. Refer to Figure 13, the smart phone includes components such as a Radio Frequency (RF) circuit 1310, a memory 1320, an input unit 1330, a display unit 1340, a sensor 1350, an audio circuit 1360, a wireless fidelity (WiFi) module 1370, a processor 1380, and a power supply 1390. The input unit 1330 may include a touch panel 1331 and other input devices 1332. The display unit 1340 may include a display panel 1341. The audio circuit 1360 may include a speaker 1361 and a microphone 1362. Those skilled in the art can understand that Figure 13 the structure of the smart phone shown in
[0210] does not limit the smart phone, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. The memory 1320 can be used to store software programs and modules. The processor 1380 runs the software programs and modules stored in the memory 1320 to execute various functional applications and data processing of the smart phone. The memory 1320 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the smart phone (such as audio data, a phone book, etc.). In addition, the memory 1320 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0211] The processor 1380 is the control center of the smart phone, connecting various parts of the entire smart phone through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 1320, and calling the data stored in the memory 1320, it executes various functions of the smart phone and processes data. Optionally, the processor 1380 may include one or more processing units; preferably, the processor 1380 can integrate an application processor and a modulation and demodulation processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modulation and demodulation processor mainly processes wireless communication. It can be understood that the above modulation and demodulation processor may not be integrated into the processor 1380.
[0212] In this embodiment, the steps executed by the processor 1380 in the device can be based on Figure 13 the structure shown in
[0213] The device may further include a server. Please refer to Figure 14 shown inFigure 14 The structural diagram of server 1400 provided by the embodiment of this application. Server 1400 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs for short) 1422 (for example, one or more processors) and a memory 1432, and one or more storage media 1430 (for example, one or more mass storage devices) for storing application programs 1442 or data 1444. Among them, the memory 1432 and the storage media 1430 may be transient storage or persistent storage. The programs stored in the storage media 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1422 may be configured to communicate with the storage media 1430 and execute a series of instruction operations in the storage media 1430 on the server 1400.
[0214] Server 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1458, and / or one or more operating systems 1441, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0215] In this embodiment, the central processing unit 1422 in the server 1400 may perform the following steps:
[0216] Obtain a first image set corresponding to the first media content and a second image set corresponding to the second media content;
[0217] Extract features from the first images in the first image set to obtain a first feature vector, and extract features from the second images in the second image set to obtain a second feature vector;
[0218] Perform object recognition on the first images in the first image set to obtain first object features, and perform object recognition on the second images in the second image set to obtain second object features;
[0219] Concatenate the first object feature and the first feature vector belonging to the same first image to obtain a first target feature vector corresponding to the first image, and concatenate the second object feature and the second feature vector belonging to the same second image to obtain a second target feature vector corresponding to the second image;
[0220] If it is determined that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector, duplicate removal processing is performed.
[0221] According to one aspect of the present application, there is provided a computer-readable storage medium for storing program code for executing the media content duplicate removal method described in each of the foregoing embodiments.
[0222] According to one aspect of the present application, there is provided a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the foregoing embodiments.
[0223] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0224] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.
[0225] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0226] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically as individual units, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0227] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0228] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A method for deduplicating media content, characterized in that, The method includes: Obtaining a first image set corresponding to a first media content and a second image set corresponding to a second media content; Performing feature extraction on the first images in the first image set to obtain first feature vectors, and performing feature extraction on the second images in the second image set to obtain second feature vectors; Performing subject recognition on the first images in the first image set to obtain first subject features, and performing subject recognition on the second images in the second image set to obtain second subject features; Concatenating the first subject features and the first feature vectors belonging to the same first image to obtain a first target feature vector corresponding to the first image, and concatenating the second subject features and the second feature vectors belonging to the same second image to obtain a second target feature vector corresponding to the second image; If it is determined that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector, perform duplicate removal processing; Wherein, a feature matching model is used to determine whether the first media content is similar to the second media content, and the training method of the feature matching model is: Obtaining a third image set corresponding to a first historical media content in a training sample and a fourth image set corresponding to a second historical media content in the training sample, and whether the first historical media content is similar to the second historical media content is identified by a target label; Determining third feature vectors corresponding to the images in the third image set and fourth feature vectors corresponding to the images in the fourth image set through a feature extraction sub-model in the feature matching model; Determining third subject features corresponding to the images in the third image set and fourth subject features corresponding to the images in the fourth image set through a subject detection sub-model in the feature matching model; Concatenating the third subject features and the third feature vectors belonging to the same image through a concatenation layer in the feature matching model to obtain a third target feature vector, and concatenating the fourth subject features and the fourth feature vectors belonging to the same image through the concatenation layer in the feature matching model to obtain a fourth target feature vector; Training the feature matching model according to the third target feature vector, the fourth target feature vector and the target label; Determining a first feature distance between adjacent video frames in the third image set or the fourth image set through a regression model; Training the regression model according to the first feature distance and a reference distance, and the training of the regression model and the training of the feature matching model are performed alternately, and the reference distance is the judgment of the distance between adjacent video frames by the model before introducing subject features.
2. The method according to claim 1, wherein The first media content and the second media content include pictures, and the first images in the first image set and the second images in the second image set are the pictures themselves.
3. The method according to claim 1, characterized in that The first media content and the second media content include videos, and the obtaining of the first image set corresponding to the first media content and the second image set corresponding to the second media content includes: Extract a plurality of first video frames from the first media content to obtain the first image set, and the first images in the first image set are arranged in the time sequence of the plurality of first video frames in the first media content; Extract a plurality of second video frames from the second media content to obtain the second image set, and the second images in the second image set are arranged in the time sequence of the plurality of second video frames in the second media content.
4. The method according to claim 3, wherein Extracting a plurality of first video frames from the first media content to obtain the first image set includes: Extract the first key video frame from the first media content; If the number of the first key video frames is less than the first preset threshold, uniformly extract the video frames before and after the first key video frame in the first media content until the total number of the extracted video frames reaches the first preset threshold to obtain the first image set; Extracting a plurality of second video frames from the second media content to obtain the second image set includes: Extract the second key video frame from the second media content; If the number of the second key video frames is less than the second preset threshold, uniformly extract the video frames before and after the second key video frame in the second media content until the total number of the extracted video frames reaches the second preset threshold to obtain the second image set.
5. The method according to claim 1, wherein Extracting the first feature vector by performing feature extraction on the first images in the first image set and extracting the second feature vector by performing feature extraction on the second images in the second image set includes: Determine the first feature vector and the second feature vector through the feature extraction sub-model in the feature matching model; Performing main body recognition on the first images in the first image set to obtain the first main body feature and performing main body recognition on the second images in the second image set to obtain the second main body feature includes: Determine the first main body feature and the second main body feature through the main body detection sub-model in the feature matching model; Concatenating the first main body feature and the first feature vector belonging to the same first image to obtain the first target feature vector corresponding to the first image, and concatenating the second main body feature and the second feature vector belonging to the same second image to obtain the second target feature vector corresponding to the second image includes: Concatenate the first main body feature and the first feature vector belonging to the same first image through the concatenation layer in the feature matching model to obtain the first target feature vector, and concatenate the second main body feature and the second feature vector belonging to the same second image through the concatenation layer in the feature matching model to obtain the second target feature vector; Determining that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector includes: Determine that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector through the matching sub-model in the feature matching model.
6. The method according to claim 1, wherein If the first historical media content and the second historical media content are videos, the images in the third image set are multiple video frames extracted from the first historical media content, and the images in the third image set are arranged in the time sequence of the video frames in the first historical media content; the images in the fourth image set are multiple video frames extracted from the second historical media content, and the images in the fourth image set are arranged in the time sequence of the video frames in the second historical media content.
7. The method according to any one of claims 1-6, characterized in that, Determining that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector includes: Determining the similarity between the first media content and the second media content according to the first target feature vector and the second target feature vector; If the similarity meets a preset condition, determining that the first media content is similar to the second media content.
8. The method according to claim 7, wherein If the first media content and the second media content include pictures, determining the similarity between the first media content and the second media content according to the first target feature vector and the second target feature vector includes: Determining a second feature distance between the first media content and the second media content according to the first target feature vector and the second target feature vector, where the second feature distance is used to represent the similarity between the first media content and the second media content; If the similarity meets a preset condition, determining that the first media content is similar to the second media content includes: If the second feature distance is less than or equal to a first distance threshold, determining that the first media content is similar to the second media content, where the preset condition is that the second feature distance is less than or equal to the first distance threshold.
9. The method according to claim 7, wherein If the first media content and the second media content include videos, determining the similarity between the first media content and the second media content according to the first target feature vector and the second target feature vector includes: Aligning the first image in the first image set with the second image in the second image set to establish a correspondence between the first image and the second image; For each pair of the first image and the second image with a correspondence, determining a third feature distance between the first image and the second image according to the first target feature vector and the second target feature vector; If the third feature distance is less than or equal to a second distance threshold, determining that the first image and the second image are similar; Obtaining the number of pairs of similar first images and second images, where the number of pairs of similar first images and second images is used to represent the similarity between the first media content and the second media content.
10. A media content deduplication device, characterized in that, The device includes an acquisition unit, an extraction unit, an identification unit, a splicing unit, a training unit, and a duplicate removal unit: The acquisition unit is configured to acquire a first image set corresponding to the first media content and a second image set corresponding to the second media content; The extraction unit is configured to extract features from the first images in the first image set to obtain first feature vectors, and extract features from the second images in the second image set to obtain second feature vectors; The recognition unit is configured to perform subject recognition on the first images in the first image set to obtain first subject features, and perform subject recognition on the second images in the second image set to obtain second subject features; The splicing unit is configured to splice the first subject features and the first feature vectors belonging to the same first image to obtain a first target feature vector corresponding to the first image, and splice the second subject features and the second feature vectors belonging to the same second image to obtain a second target feature vector corresponding to the second image; The duplicate removal unit is configured to perform duplicate removal processing if it is determined that the first media content is similar to the second media content according to the first target feature vector and the second target feature vector; Wherein, a feature matching model is used to determine whether the first media content is similar to the second media content; The training unit is configured to: Obtain a third image set corresponding to the first historical media content in the training sample, and a fourth image set corresponding to the second historical media content in the training sample, whether the first historical media content and the second historical media content are similar is identified by a target label; Determine third feature vectors corresponding to the images in the third image set and fourth feature vectors corresponding to the images in the fourth image set through a feature extraction sub-model in the feature matching model; Determine third subject features corresponding to the images in the third image set and fourth subject features corresponding to the images in the fourth image set through a subject detection sub-model in the feature matching model; Splice the third subject features and the third feature vectors belonging to the same image through a splicing layer in the feature matching model to obtain a third target feature vector, and splice the fourth subject features and the fourth feature vectors belonging to the same image through the splicing layer in the feature matching model to obtain a fourth target feature vector; Train the feature matching model according to the third target feature vector, the fourth target feature vector and the target label; Determine a first feature distance between adjacent video frames in the third image set or the fourth image set through a regression model; Train the regression model according to the first feature distance and a reference distance, the training of the regression model and the training of the feature matching model are performed alternately, and the reference distance is the judgment of the distance between adjacent video frames by the model before introducing the subject features.
11. The device according to claim 10, characterized in that, The first media content and the second media content include pictures, and the first images in the first image set and the second images in the second image set are the pictures themselves.
12. The device according to claim 10, wherein, The first media content and the second media content include videos, and the acquisition unit is configured to: Extract a plurality of first video frames from the first media content to obtain the first image set, and the first images in the first image set are arranged in the time sequence of the plurality of first video frames in the first media content; Extract a plurality of second video frames from the second media content to obtain the second image set, and the second images in the second image set are arranged in the time sequence of the plurality of second video frames in the second media content; the number of first video frames in the first image set is the same as the number of second video frames in the second image set.
13. The device according to claim 12, characterized in that, The obtaining unit is specifically configured to: Extract first key video frames from the first media content; If the number of the first key video frames is less than a first preset threshold, uniformly extract the video frames before and after the first key video frames in the first media content until the total number of the extracted video frames reaches the first preset threshold to obtain the first image set; Extract second key video frames from the second media content; If the number of the second key video frames is less than a second preset threshold, uniformly extract the video frames before and after the second key video frames in the second media content until the total number of the extracted video frames reaches the second preset threshold to obtain the second image set.
14. The device according to any one of claims 10 to 13, characterized in that, The duplicate removal unit is specifically configured to: Determine the similarity between the first media content and the second media content according to the first target feature vector and the second target feature vector; If the similarity meets a preset condition, determine that the first media content is similar to the second media content.
15. The device according to claim 14, characterized in that, If the first media content and the second media content include pictures, the duplicate removal unit is specifically configured to: Determine a second feature distance between the first media content and the second media content according to the first target feature vector and the second target feature vector, and the second feature distance is used to represent the similarity between the first media content and the second media content; If the similarity meets a preset condition, determining that the first media content is similar to the second media content includes: If the second feature distance is less than or equal to a first distance threshold, determine that the first media content is similar to the second media content, and the preset condition is that the second feature distance is less than or equal to the first distance threshold.
16. The device according to claim 14, wherein If the first media content and the second media content include videos, the duplicate removal unit is configured to: Align the first images in the first image set with the second images in the second image set to establish a corresponding relationship between the first images and the second images; For each pair of corresponding first image and second image, determine a third feature distance between the first image and the second image according to the first target feature vector and the second target feature vector; If the third feature distance is less than or equal to a second distance threshold, determine that the first image and the second image are similar; Obtain the number of pairs of similar first images and second images, and the number of pairs of similar first images and second images is used to represent the similarity between the first media content and the second media content.
17. A device for media content deduplication, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the method according to any one of claims 1-9 based on the instructions in the program code.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1-9.
19. A computer program product, characterized in that, The computer program product includes computer instructions, and the processor of the computer device executes the computer instructions, so that the computer device executes the method according to any one of claims 1-9.
Citation Information
Patent Citations
Video auditing method and device and electronic equipment
CN110225373A
Short video copyright detection method and system
CN111182364A
Similar video processing method and device based on artificial intelligence and electronic equipment
CN112203122A