Visual Media Data Deduplication Processing Method, Apparatus, Device, and Storage Medium

By combining visual features and text content features of visual media data, the problem of low deduplication rate in the prior art is solved, and more efficient deduplication processing of visual media data is achieved.

CN114329050BActive Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111541971.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-07-04
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The existing visual media data deduplication method based on md5 value is sensitive to operations such as compression, cropping and watermarking in visual media data, resulting in a low deduplication rate and the inability to effectively identify the same but mildly edited visual media data.

Method used

By extracting visual features of visual media data, including image features and text area features, and performing similarity analysis in combination with text content features, deduplication of visual media data is achieved.

Benefits of technology

It improves the deduplication rate and deduplication processing efficiency of visual media platforms, avoids excessive attention to interference factors, improves the accuracy of similarity calculations, and reduces the omissions of repeated videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114329050B_ABST
    Figure CN114329050B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, device, and storage medium for deduplicating visual media data. The method involves artificial intelligence and includes: respectively extracting visual features from at least two visual media data to obtain the visual features of each visual media data, where the visual features include image features and text region features. By extracting text information from each visual media data, the text content features of each visual media data are obtained, and then, based on the visual features and the text content features, similarity analysis is performed on at least two visual media data to obtain the similarity between the visual media data, and deduplication processing is performed on at least two visual media data according to the similarity between the visual media data. By considering from multiple angles using this method, the accuracy of the calculated similarity between videos can be improved, the situation of missing duplicate videos without deduplication can be avoided, and the video deduplication rate and deduplication processing efficiency of the video platform are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a method, apparatus, device, and storage medium for deduplicating visual media data. Background Art

[0002] With the development of computer technology and the wide application of the Internet in people's lives and work, more and more people obtain and transmit information through the Internet. Among them, the information forms can be text, web pages, audio data, and visual media data, etc. As the data form with the most comprehensive content, visual media data has become the information transmission method of most users.

[0003] Due to the low threshold of existing visual media production methods, users can quickly generate visual media with the help of visual media production tools, and thus a large amount of visual media is published on the network at all times. However, there are situations where the visual media data published by different users is repeated or copied from others, which easily leads to the problem that a large amount of visual media data is repeated, occupying the publishing channel and the display interface resources. Therefore, for a visual media platform, it is necessary to pay attention to the visual media data published on the platform in real time and perform deduplication processing on the visual media data to improve the quality of the published visual media data and attract more users.

[0004] Traditionally, the deduplication method based on the md5 value is mostly used. That is, first calculate the md5 value (i.e., MD5 information digest value) of the visual media data, and then determine whether the visual media data is the same according to the md5 value, and further perform deduplication on the visual media data. However, the traditional deduplication method based on the md5 value is relatively sensitive to the interference factors in the visual media data, such as compression, cropping, watermarking, etc. during the upload process of the visual media data. That is to say, after the same visual media data is compressed or mildly edited (operations such as cropping and watermarking), it is easily recognized as non-repeated visual media data. Therefore, the deduplication rate of the traditional deduplication method based on the md5 value still needs to be improved. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, device, and storage medium for deduplicating visual media data that can improve the deduplication rate of visual media data on a visual media platform.

[0006] A method for deduplicating visual media data, the method comprising:

[0007] Extract visual features from at least two visual media data respectively to obtain the visual features of each visual media data, where the visual features include image features and text region features;

[0008] Extract text information from each of the visual media data to obtain the text content features of each of the visual media data;

[0009] Based on the visual features and the text content features, perform a similarity analysis on the at least two visual media data to obtain the similarity between the visual media data;

[0010] According to the similarity between the visual media data, perform a duplicate removal process on the at least two visual media data.

[0011] A visual media data device, the device includes:

[0012] A visual feature extraction module, configured to respectively extract visual features from at least two visual media data to obtain the visual features of each of the visual media data, where the visual features include image features and text region features;

[0013] A text content feature extraction module, configured to extract text information from each of the visual media data to obtain the text content features of each of the visual media data;

[0014] A similarity analysis module, configured to perform a similarity analysis on the at least two visual media data based on the visual features and the text content features to obtain the similarity between the visual media data;

[0015] A duplicate removal processing module, configured to perform a duplicate removal process on the at least two visual media data according to the similarity between the visual media data.

[0016] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0017] Respectively extract visual features from at least two visual media data to obtain the visual features of each of the visual media data, where the visual features include image features and text region features;

[0018] Extract text information from each of the visual media data to obtain the text content features of each of the visual media data;

[0019] Based on the visual features and the text content features, perform a similarity analysis on the at least two visual media data to obtain the similarity between the visual media data;

[0020] According to the similarity between the visual media data, perform a duplicate removal process on the at least two visual media data.

[0021] A computer-readable storage medium storing a computer program, which when executed by a processor implements the following steps:

[0022] Extract visual features from at least two visual media data respectively to obtain the visual features of each visual media data, where the visual features include image features and text region features;

[0023] Extract text information from each visual media data to obtain the text content features of each visual media data;

[0024] Based on the visual features and the text content features, perform a similarity analysis on the at least two visual media data to obtain the similarity between the visual media data;

[0025] According to the similarity between the visual media data, perform a deduplication process on the at least two visual media data.

[0026] A computer program product including a computer program, which when executed by a processor implements the following steps:

[0027] Extract visual features from at least two visual media data respectively to obtain the visual features of each visual media data, where the visual features include image features and text region features;

[0028] Extract text information from each visual media data to obtain the text content features of each visual media data;

[0029] Based on the visual features and the text content features, perform a similarity analysis on the at least two visual media data to obtain the similarity between the visual media data;

[0030] According to the similarity between the visual media data, perform a deduplication process on the at least two visual media data.

[0031] In the above method, apparatus, device, and storage medium for deduplicating visual media data, visual features of at least two visual media data are extracted respectively to obtain the visual features of each visual media data. Among them, the visual features include image features and text region features. By considering the text region features on the visual media data simultaneously, the situation of classifying visual media data with the same image features and different text region features into the same visual media data is avoided. By extracting text information from each visual media data, the text content features of each visual media data are obtained. Furthermore, based on the visual features and the text content features, similarity analysis can be performed on at least two visual media data to obtain the similarity between the visual media data, and deduplication processing can be performed on at least two visual media data according to the similarity between the visual media data. It realizes comprehensive consideration from multiple perspectives to improve the accuracy of the similarity degree calculated between videos, avoids the situation of omitting duplicate videos without deduplication, and at the same time, by adopting the method of comprehensive consideration from multiple perspectives, it can also avoid the problem of over-concern about interference factors in the traditional deduplication method based on md5 values, improving the video deduplication rate and deduplication processing efficiency of the video platform. Description of the Drawings

[0032] Figure 1 It is an application environment diagram of the method for deduplicating visual media data in an embodiment;

[0033] Figure 2 It is a schematic flowchart of the method for deduplicating visual media data in an embodiment;

[0034] Figure 3 It is a schematic flowchart of the process of deduplicating at least two visual media data in an embodiment;

[0035] Figure 4 It is a schematic flowchart of the process of obtaining a trained feature extraction network in an embodiment;

[0036] Figure 5 It is a schematic flowchart of the method for deduplicating visual media data in another embodiment;

[0037] Figure 6 It is a schematic flowchart of the method for deduplicating visual media data in still another embodiment;

[0038] Figure 7 It is a structural block diagram of the apparatus for deduplicating visual media data in an embodiment;

[0039] Figure 8 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0040] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0041] The video duplicate removal processing method provided by this application involves artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theory, method, technology and application system. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, machine learning / deep learning, autonomous driving, and intelligent transportation.

[0042] Among them, computer vision technology (CV) Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking and measurement, and further perform graphics processing to make the computer process into an image that is more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0043] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0044] The video duplicate removal processing method provided in this application involves technologies such as computer vision in artificial intelligence and can be applied to the application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. The server 104 extracts visual features from at least two visual media data respectively to obtain the visual features of each visual media data, and extracts text information from each visual media data to obtain the text content features of each visual media data. Among them, the visual features include image features and text region features. The server 104 can receive the local visual media data sent by the terminal 102, or can obtain the visual media data from the cloud database of the server 104 itself. Further, the server 104 can perform a similarity analysis on at least two visual media data based on the visual features and the text content features to obtain the similarity between the visual media data, and then, according to the similarity between the visual media data, perform duplicate removal processing on at least two visual media data to obtain the duplicate-removed visual media data and save it. Among them, the server 104 can save the duplicate-removed visual media data in its own cloud database, or can send the duplicate-removed visual media data to the terminal 102 for display and storage. Among them, the terminal 102 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle terminal, a smart TV, etc., but is not limited thereto. The server 104 can be an independent physical server, or can be a server cluster or a distributed system composed of multiple physical servers, or can also be a cloud server providing cloud computing services. The terminal 102 and the server 104 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0045] In one embodiment, as Figure 2 shown, a method for removing duplicates from visual media data is provided. Taking the server in Figure 1 as an example, the method includes the following steps:

[0046] Step S202: Extract visual features from at least two visual media data respectively to obtain the visual features of each visual media data. The visual features include image features and text region features.

[0047] Specifically, obtain a visual media database to be de-duplicated. Among them, the visual media database to be de-duplicated includes at least two visual media data to be de-duplicated. Then, extract visual features from at least two visual media data in the visual media database to obtain the visual features of the visual media data.

[0048] Among them, the visual features of the visual media data specifically include image features and text region features. The image features represent the image data represented by the visual media data, while the text region features represent the specific regions where the text is located in the visual media data with text.

[0049] Furthermore, the visual media data can include video data and image data. The video data can include video data of different durations or different video platforms. The image data can include different types and uses of image data such as pictures and emoticons.

[0050] In one embodiment, taking the visual media data as video data as an example, before extracting visual features from at least two visual media data respectively to obtain the visual features of each visual media data, it further includes: extracting video frames from at least two video data to obtain video frames.

[0051] Specifically, the purpose of video frame extraction is to select representative image frames in the video for measuring video similarity to reduce the computational complexity in the similarity analysis process. Among them, various methods can be used for video frame extraction, such as fixed-interval frame extraction and key-frame extraction.

[0052] Among them, key-frame extraction can extract specific key frames of interest from the video according to the actual situation. For example, object detection and recognition algorithms can be used to extract image frames containing the objects of interest in the video. Specifically, the target video information, that is, the video information of interest to the user, can be obtained, and based on the target video information, key-frame extraction is performed on at least two video data to obtain video key frames.

[0053] Similarly, fixed-interval frame extraction means extracting one frame at a fixed time interval. Specifically, it can be to perform timed video frame extraction on at least two video data according to a preset frame extraction time interval to obtain corresponding fixed video frames. Among them, the preset frame extraction time interval can be adjusted and modified according to user needs or actual application scenarios, and no specific limitation is made.

[0054] Furthermore, when extracting frames, it is necessary to ensure that the frame extraction algorithm has no randomness, that is, for the same video, the frames extracted each time should be the same, so as to ensure that the frames extracted from two identical videos are also the same, which is convenient for calculating similarity. Among them, the number of frames extracted from each video should be fixed, such as 10 frames, and the excess frames are discarded, and the insufficient frames can be filled with the last frame.

[0055] Step S204: Extract text information from each visual media data to obtain the text content features of each visual media data.

[0056] Specifically, through image preprocessing of each visual media data, the preprocessed image frame to be recognized is obtained, and based on the image frame to be recognized, character segmentation and character recognition are performed to obtain the segmented characters in sequence. Furthermore, through dimensionality reduction processing and feature extraction of the segmented characters, character features are obtained. Further, feature classification and content recognition can be performed based on the character features to obtain the text content features of the visual media data.

[0057] Among them, through image preprocessing of each visual media data, such as grayscale processing or binary processing, as well as noise reduction processing and other image preprocessing operations, the preprocessed image frame to be recognized is obtained. Then, through character segmentation and character recognition of the image frame to be recognized, where character segmentation is to split the characters in the image frame to be recognized into individual characters, and character recognition is performed on the segmented characters in sequence. Among them, if the character line is inclined, it is necessary to further correct the inclination of the characters, and further perform normalization processing on the segmented characters, that is, adjust the individual character images to the same size specification, and then perform character recognition on the segmented characters one by one.

[0058] Furthermore, through dimensionality reduction processing and feature extraction of the segmented characters, character features are obtained. Among them, characters can include different character types such as numbers, letters, symbols, and Chinese characters. By performing feature extraction on the characters, the character types can be further determined. Among them, for numbers, letters, or symbols, since the number of numbers, letters, or symbols themselves is small and can be classified as a small character set, dimensionality reduction processing may not be performed or simple dimensionality reduction processing can be performed to meet the requirements of character recognition.

[0059] For Chinese characters, since the number of Chinese characters is large, it belongs to a large character set, and the structure of Chinese characters is complex and there are many similar-shaped characters, it is difficult to perform character recognition. Therefore, in order to improve the efficiency of character recognition, it is necessary to perform dimensionality reduction processing on Chinese characters to reduce the feature dimension. At the same time, in order to ensure that the feature vector after reducing the dimension retains sufficient character information to achieve the purpose of distinguishing different characters, the intensity of the dimensionality reduction processing needs to be adjusted and modified according to actual needs.

[0060] In one embodiment, after dimension reduction processing and feature extraction are performed on the segmented characters, and character features are obtained, feature classification and content recognition are further performed based on the character features to obtain the text content features of the visual media data.

[0061] Specifically, by using a trained classifier to perform feature classification and content recognition on the character features, it is determined which character classification each character feature belongs to. Furthermore, after determining the character classification of the character features, the character content can be further obtained to obtain the text content features of the visual media data.

[0062] Among them, a character library can be specifically obtained by random collection, and the initial classifier can be trained according to the character library to obtain a trained classifier.

[0063] In one embodiment, relevant algorithms or network models for text information extraction can be used to extract text information from the visual media data to obtain the text content features of each visual media data. For example, the OCR (Optical Character Recognition) algorithm is used to extract text information from the visual media data. In this embodiment, the relevant algorithms or network models for text information extraction used are not specifically limited, and the algorithms or network models used can meet the requirements of text information extraction.

[0064] Step S206: Based on the visual features and the text content features, perform a similarity analysis on at least two visual media data to obtain the similarity between the visual media data.

[0065] Specifically, based on the image features and the text region features, visual similarity calculation is performed to generate the visual similarity values of at least two visual media data, and based on the text content features, text similarity calculation is performed to generate the text similarity values of at least two visual media data. Furthermore, by combining the visual similarity values and the text similarity values, the similarity between the visual media data is obtained.

[0066] Among them, the Euclidean distance can be used to calculate the visual similarity of two visual media data, that is, based on the image features and the text region features, the Euclidean distance between the two visual media data is calculated, and according to the calculated Euclidean distance value, the visual similarity degree of the two visual media data is judged. Among them, the smaller the Euclidean distance, the higher the visual similarity degree of the two visual media data.

[0067] In addition, when the visual media data is a picture, in this embodiment, the perceptual hash algorithm (i.e., Perceptual hash algorithm) can also be used to calculate the visual similarity of two visual media data. Among them, the role of the perceptual hash algorithm is to generate a fingerprint string for each picture, and then compare the fingerprint strings of different pictures, and determine the visual similarity degree of the two visual media data according to the comparison result of the fingerprint strings. The closer the fingerprint strings are, the more similar the pictures are, that is, the higher the visual similarity degree of the two visual media data.

[0068] In one embodiment, the text similarity can be calculated according to the text content features and the edit distance corresponding to the text content, and the text similarity values of at least two visual media data are generated.

[0069] Among them, the edit distance refers to the minimum number of steps required to change one piece of text into another piece of text through three operations: deletion, addition, and replacement. It can be understood that the smaller the edit distance corresponding to the text content of two pieces of text, the higher the similarity degree of the two pieces of text.

[0070] In addition, due to the high computational complexity of the edit distance calculation, in scenarios with high timeliness requirements, the faster Jaccard similarity calculation can also be used. Among them, the Jaccard similarity (i.e., Jaccard similarity coefficient) is mainly used to compare the similarity and difference between finite sample sets, and can calculate the similarity between samples with symbolic metrics or boolean value metrics. For two pieces of text, the Jaccard similarity is the number of elements in the intersection of the two pieces of text divided by the number of elements in the union.

[0071] Furthermore, by calculating the sum of the visual similarity value and the text similarity value, the similarity between visual media data can be obtained.

[0072] Step S208, perform duplicate removal processing on at least two visual media data according to the similarity between the visual media data.

[0073] Specifically, by obtaining a preset similarity threshold and comparing the similarity between the visual media data with the preset similarity threshold, it is determined whether the similarity between the visual media data is greater than the preset similarity threshold. When it is determined that the similarity between the visual media data is greater than the preset similarity threshold, it indicates that there are duplicate data in the two currently compared visual media data, and then duplicate removal processing needs to be performed to retain the visual media data after duplicate removal.

[0074] In one embodiment, if the visual media data includes video data, before performing visual feature extraction, it is also necessary to extract frames from at least two video data to obtain video frames, and then perform visual feature extraction on the video frames to obtain the visual features of each video frame, including image features and text region features, and perform text information extraction on each video frame to obtain the text content features of each video frame.

[0075] Specifically, based on the visual features and the text content features, perform similarity analysis on at least two video frames to obtain the similarity between two video frames, and obtain a preset similarity threshold, and then determine whether the similarity between the video frames is greater than the preset similarity threshold.

[0076] Among them, when it is determined that the similarity between the video frames is greater than the preset similarity threshold, determine that the current two video frames are a pair of similar video frames, and obtain the number of pairs of similar video frames in the preset pair of video frames extracted from any two video data.

[0077] Furthermore, by obtaining a preset threshold for pairs of similar video frames and based on the preset threshold for pairs of similar video frames and the number of pairs of similar video frames, determine whether there are duplicate videos in the current two video data. When it is determined that there are duplicate videos, perform video deduplication processing on the current two video data.

[0078] In the above method for deduplicating visual media data, by separately performing visual feature extraction on at least two visual media data to obtain the visual features of each visual media data, where the visual features include image features and text region features, by simultaneously considering the text region features on the visual media data, it is possible to avoid the situation where visual media data with the same image features and different text region features are classified as the same visual media data. By performing text information extraction on each visual media data to obtain the text content features of each visual media data, it is then possible to perform similarity analysis on at least two visual media data based on the visual features and the text content features to obtain the similarity between the visual media data, and perform deduplication processing on at least two visual media data according to the similarity between the visual media data. This realizes comprehensive consideration from multiple perspectives to improve the accuracy of the calculated similarity between videos, avoid the situation of missing duplicate videos without deduplication, and at the same time, by adopting the method of comprehensive consideration from multiple perspectives, it is also possible to avoid the problem of over - attention to interference factors in traditional deduplication methods based on md5 values, improving the video deduplication rate and deduplication processing efficiency of the video platform.

[0079] In one embodiment, the step of separately performing visual feature extraction on at least two visual media data to obtain the visual features of each visual media data specifically includes:

[0080] When text is detected in visual media data, based on the trained feature extraction network, the text region positions of at least two visual media data are detected respectively to generate the text region features of the visual media data;

[0081] Based on the trained feature extraction network, the image features of at least two visual media data are extracted respectively to obtain the image features of at least two visual media data.

[0082] Specifically, in visual media data, the text information therein is crucial. The same visual picture with different texts will cause a large change in the meaning of the visual media data. Therefore, when extracting visual features, it is necessary to consider the position of the text in the picture at the same time, that is, the text region position.

[0083] In this embodiment, specifically, based on the trained feature extraction network, the text region position in the visual media data is detected to generate the text region feature. Among them, the text region feature can specifically be a text region mask, and the text region mask is a matrix with values of 0 or 1, where 1 represents the text region and 0 represents the non-text region.

[0084] Further, after obtaining the text region mask, the text region mask is used as a channel to form a four-channel with the RGB channels of the image as the input data of the trained feature extraction network. It can be understood that when the trained feature extraction network extracts features from the input data composed of the text region mask and the RGB channel data of the image, the text region features and image features of the visual media data can be output.

[0085] In this embodiment, when text is detected in visual media data, based on the trained feature extraction network, the text region positions of at least two visual media data are detected respectively to generate the text region features of the visual media data, and based on the trained feature extraction network, the image features of at least two visual media data are extracted respectively to obtain the image features of at least two visual media data. When extracting visual features from visual media data, it is possible to consider both the visual picture of the data media data and the text region position on the visual picture at the same time, so as to avoid classifying visual media data with the same image features and different text region features as the same visual media data, thereby improving the subsequent deduplication rate of visual media data, reducing duplicate deduplication processing operations, and further improving the deduplication processing efficiency.

[0086] In one embodiment, as Figure 3As shown, when the visual media data is video data and the similarity between visual media data is the similarity between video frames, the step of deduplicating at least two visual media data, that is, the step of deduplicating at least two visual media data according to the similarity between visual media data, specifically includes:

[0087] Step S302, obtain a preset similarity threshold, and determine whether the similarity between video frames is greater than the preset similarity threshold.

[0088] Specifically, by obtaining a preset similarity threshold and comparing the preset similarity threshold with the similarity between video frames, it is determined whether the similarity between video frames is greater than the preset similarity threshold. Among them, the preset similarity threshold can be adjusted and modified according to user requirements and actual application scenarios, and is not limited to certain specific values.

[0089] Step S304, when it is determined that the similarity between video frames is greater than the preset similarity threshold, determine the current two video frames as a pair of similar video frames.

[0090] Specifically, by comparing the preset similarity threshold with the similarity between video frames, and when it is determined that the similarity between video frames is greater than the preset similarity threshold, it indicates that the two currently compared video frames are a pair of similar video frames. Among them, when the similarity between video frames is not greater than the preset similarity threshold, it indicates that the two currently compared video frames do not belong to similar video frames.

[0091] Step S306, obtain the number of pairs of similar video frames in the preset pairs of video frames extracted from any two video data.

[0092] Specifically, for the video frames extracted from any two video data, similarity analysis is respectively performed to determine all pairs of similar video frames, and the number of pairs of similar video frames is counted.

[0093] In one embodiment, assume that 10 frames are extracted from each video. For videos A and B, first calculate the similarity between the 10 video frames extracted from video A and the 10 video frames extracted from video B, and determine whether the similarity between the 10 video frames extracted respectively is greater than the preset similarity threshold. Among them, the number of frames extracted from the video is not specifically limited either, and can be modified and adjusted according to actual needs and different application scenarios.

[0094] Specifically, for each frame extracted from video A, the video frame with the highest similarity is sequentially matched from video B to obtain 10 pairs of video frames, and the number of pairs of similar video frames in the 10 pairs of video frames is counted. Among them, the pairs of similar video frames in the 10 pairs of video frames refer to the pairs of video frames with a similarity greater than the preset similarity threshold.

[0095] Step S308: Determine whether there are duplicate videos in the current two video data according to the preset threshold of similar video frame pairs and the number of similar video frame pairs.

[0096] Specifically, by counting the number of similar video frame pairs between the two video data, obtaining the preset threshold of similar video frame pairs, and comparing the counted number of similar video frame pairs with the preset threshold of similar video frame pairs to determine whether the number of similar video frame pairs is greater than the preset threshold of similar video frame pairs.

[0097] Among them, when the number of similar video frame pairs is greater than the preset threshold of similar video frame pairs, it is determined that there are duplicate videos in the current two video data.

[0098] For example, for each frame extracted from Video A, sequentially match the video frame with the highest similarity in Video B to obtain 10 pairs of video frames, and count the number of similar video frame pairs among the 10 pairs of video frames. If the number of similar video frame pairs is greater than the preset threshold of similar video frame pairs (for example, the preset threshold of similar video frame pairs is 9 pairs), it is determined that the videos are similar and there are duplicate videos in the current two video data.

[0099] Among them, the preset threshold of similar video frame pairs can be adjusted and modified according to user requirements or actual application scenarios, not limited to specific values. Its limiting condition is that the preset threshold of similar video frame pairs needs to be less than or equal to the number of video frame pairs extracted from any two videos.

[0100] Step S310: When it is determined that there are duplicate videos, perform video deduplication processing on the current two video data.

[0101] Specifically, when it is determined that the number of similar video frame pairs is greater than the preset threshold of similar video frame pairs, it is determined that there are duplicate videos in the two currently compared videos, and video deduplication processing is performed on the current two video data, that is, one of the videos is deleted, and one of the videos is retained and stored.

[0102] Among them, when multiple videos need to be deduplicated, the method of randomly selecting two videos for comparison and retaining one of them can be adopted. The retained video is continued to be compared with any one of the multiple videos to be compared until there are no duplicate videos at the end, then the current deduplication operation is completed.

[0103] In this embodiment, by obtaining a preset similarity threshold and determining whether the similarity between video frames is greater than the preset similarity threshold, when it is determined that the similarity between video frames is greater than the preset similarity threshold, the current two video frames are determined as a pair of similar video frames. Further, obtain the number of pairs of similar video frames in the preset pairs of video frames extracted from any two video data, and determine whether there are duplicate videos in the current two video data according to the preset threshold of pairs of similar video frames and the number of pairs of similar video frames. When it is determined that there are duplicate videos, perform video deduplication processing on the current two video data. It realizes calculating the similarity starting from the video frames of the video data, determining the corresponding pairs of similar video frames, and after determining all pairs of similar video frames, further determining whether the two videos being compared are the same based on the number of pairs of similar video frames, improving the accuracy of calculating and comparing the similarity of videos, avoiding the situation of missing duplicate videos without deduplication, and improving the video deduplication processing efficiency of the video platform.

[0104] In one embodiment, as Figure 4 shown, the steps of obtaining the trained feature extraction network specifically include:

[0105] Step S402, randomly collect a set of visual media data samples.

[0106] Specifically, by randomly collecting a set of visual media data samples, such as video data, picture data, and emoji data, etc.

[0107] Step S404, obtain different regional images of the visual media data in the set of visual media data samples as positive samples for training visual media data, and the positive samples for training visual media data include regional images with text region features.

[0108] Specifically, for the same visual media data in the set of visual media data samples, different regional images of the visual media data need to be obtained as positive samples for training visual media data. Among them, the purpose of obtaining images of different regions of the same visual media data is to determine whether the visual media data contains text and the specific position area of the text in the visual media data. Furthermore, the positive samples for training visual media data include regional images with text region features.

[0109] Step S406, obtain the same regional images of different visual media data in the set of visual media data samples as negative samples for training visual media data.

[0110] Specifically, for different visual media data in the set of visual media data samples, images of the same regions of different visual media data need to be obtained. For example, images of one region of some current visual media data, or images of multiple corresponding regions of some visual media data can also be obtained simultaneously as negative samples for training visual media data.

[0111] Step S408: Train the initial feature extraction network based on the positive training visual media data samples and the negative training visual media data samples to obtain a trained feature extraction network.

[0112] Specifically, based on the positive training visual media data samples and the negative training visual media data samples, a training visual media data sample set can be obtained. Then, the initial feature extraction network is trained according to the training visual media data sample set to obtain a trained feature extraction network.

[0113] Among them, in order to reduce the annotation amount, a self-supervised learning method can be used for training, that is, different regions of the same image are randomly used as positive samples, and different parts of different images are used as negative samples to train the similarity network. Finally, a trained feature extraction network is obtained to extract the image features and text region features of video frames. Among them, the initial feature extraction network to be trained can be network models of various different types or structures, as long as it can achieve feature extraction to meet the requirements, and its specific type is not limited.

[0114] In this embodiment, by randomly collecting a visual media data sample set and obtaining different region images of the visual media data in the visual media data sample set as positive training visual media data samples, and the region images including text region features are included in the positive training visual media data samples. By obtaining the same region images of different visual media data in the visual media data sample set as negative training visual media data samples, and then training the initial feature extraction network according to the positive training visual media data samples and the negative training visual media data samples, a trained feature extraction network is obtained. It realizes the training of the initial feature extraction network according to the negative training visual media data samples and the positive training visual media data samples including region images with text region features, which can enable the trained feature extraction network to simultaneously extract the image features and text region features of visual media data, avoid the situation of classifying visual media data with the same image features and different text region features into the same visual media data, so as to improve the subsequent deduplication rate of visual media data, reduce the repeated deduplication processing operations, and further improve the deduplication processing efficiency of visual media data.

[0115] In one embodiment, as Figure 5 shown, a method for deduplicating visual media data is provided, and the method specifically includes the following steps:

[0116] Step S501: Randomly collect a visual media data sample set.

[0117] Step S502: Obtain the images of different regions of the visual media data in the visual media data sample set as the positive samples for training visual media data. The positive samples for training visual media data include the region images with text region features.

[0118] Step S503: Obtain the images of the same region of different visual media data in the visual media data sample set as the negative samples for training visual media data.

[0119] Step S504: Train the initial feature extraction network according to the positive samples and negative samples for training visual media data to obtain a trained feature extraction network.

[0120] Step S505: When it is detected that there is text in the visual media data, based on the trained feature extraction network, detect the positions of text regions for at least two visual media data respectively to generate the text region features of the visual media data.

[0121] Step S506: Based on the trained feature extraction network, extract image features for at least two visual media data respectively to obtain the image features of at least two visual media data.

[0122] Step S507: Perform image preprocessing on each visual media data to obtain a preprocessed image frame to be recognized.

[0123] Step S508: Based on the image frame to be recognized, perform character segmentation and character recognition to obtain the segmented characters in sequence.

[0124] Step S509: Perform dimensionality reduction processing and feature extraction on the segmented characters to obtain character features.

[0125] Step S510: Based on the character features, perform feature classification and content recognition to obtain the text content features of the visual media data.

[0126] Step S511: Based on the image features and text region features, perform visual similarity calculation to generate the visual similarity values of at least two visual media data.

[0127] Step S512: Based on the text content features, perform text similarity calculation to generate the text similarity values of at least two visual media data.

[0128] Step S513: Combine the visual similarity values and text similarity values to obtain the similarity between visual media data.

[0129] Step S514: Based on the similarity between visual media data, perform duplicate removal processing on at least two visual media data.

[0130] In the above method for deduplicating visual media data, by simultaneously considering the text region features on the visual media data, it is avoided that visual media data with the same image features but different text region features are classified as the same visual media data. Considering from multiple perspectives in combination can improve the accuracy of the calculated similarity degree between videos, avoid the situation of missing duplicate videos without deduplication, and at the same time, by adopting the method of comprehensive consideration from multiple perspectives, it can also avoid the problem of over - attention to interference factors in traditional deduplication methods based on md5 values, improving the video deduplication rate and deduplication processing efficiency of the video platform.

[0131] In one embodiment, as Figure 6 shown, a method for deduplicating visual media data is provided. When the visual media data is video data, the method specifically includes the following steps:

[0132] Step S601, perform video frame extraction on at least two video data to obtain video frames.

[0133] Step S602, when it is detected that there is text in the video frame, based on the trained feature extraction network, respectively perform text region position detection on at least two video frames to generate text region features of the video frames.

[0134] Step S603, based on the trained feature extraction network, respectively extract image features from at least two video frames to obtain image features of at least two video frames.

[0135] Step S604, perform image pre - processing on each video frame to obtain the pre - processed image frame to be recognized.

[0136] Step S605, based on the image frame to be recognized, perform character segmentation and character recognition to sequentially obtain the segmented characters.

[0137] Step S606, perform dimensionality reduction processing and feature extraction on the segmented characters to obtain character features.

[0138] Step S607, based on the character features, perform feature classification and content recognition to obtain the text content features of the video frame.

[0139] Step S608, based on the image features and text region features, perform visual similarity calculation to generate visual similarity values of at least two video frames.

[0140] Step S609, based on the text content features, perform text similarity calculation to generate text similarity values of at least two video frames.

[0141] Step S610, comprehensively combine the visual similarity values and text similarity values to obtain the similarity between video frames.

[0142] Step S611, obtain a preset similarity threshold, and determine whether the similarity between video frames is greater than the preset similarity threshold.

[0143] Step S612, when it is determined that the similarity between video frames is greater than the preset similarity threshold, determine the current two video frames as a pair of similar video frames.

[0144] Step S613, obtain the number of pairs of similar video frames in the preset pairs of video frames extracted from any two video data.

[0145] Step S614, determine whether there are duplicate videos in the current two video data according to the preset threshold of pairs of similar video frames and the number of pairs of similar video frames.

[0146] Step S615, when it is determined that there are duplicate videos, perform video deduplication processing on the current two video data.

[0147] In the above method for deduplicating visual media data, by simultaneously considering the text region features on the visual media data, it is avoided that visual media data with the same image features and different text region features are classified as the same visual media data, and a combination consideration is carried out from multiple perspectives to improve the accuracy of the similarity degree calculated between videos, avoid the situation where duplicate videos are missed without deduplication, and at the same time, by adopting the method of comprehensive consideration from multiple perspectives, it is also possible to avoid the problem of excessive attention to interference factors in the traditional deduplication method based on md5 values, improving the video deduplication rate and deduplication processing efficiency of the video platform.

[0148] This application also provides an application scenario, which applies the above method for deduplicating visual media data. Specifically, the application of the method for deduplicating visual media data in this application scenario is as follows:

[0149] When the visual media data is a picture, directly perform visual feature extraction on at least two pictures respectively to obtain the visual features of each picture, where the visual features include image features and text region features. At the same time, it is also necessary to perform text information extraction on each picture to obtain the text content features of each picture, and then, based on the visual features and the text content features, perform similarity analysis on at least two visual media data to obtain the similarity between pictures, so as to perform deduplication processing on at least two pictures according to the similarity between pictures.

[0150] It should be understood that although the steps in the respective flowcharts involved in the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the respective flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0151] In one embodiment, as Figure 7 shown, a visual media data deduplication processing device is provided. This device can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the device includes: a visual feature extraction module 702, a text content feature extraction module 704, a similarity analysis module 706, and a deduplication processing module 708, where:

[0152] The visual feature extraction module 702 is used to extract visual features from at least two visual media data respectively to obtain the visual features of each visual media data. The visual features include image features and text region features.

[0153] The text content feature extraction module 704 is used to extract text information from each visual media data to obtain the text content features of each visual media data.

[0154] The similarity analysis module 706 is used to perform similarity analysis on at least two visual media data based on the visual features and the text content features to obtain the similarity between the visual media data.

[0155] The deduplication processing module 708 is used to perform deduplication processing on at least two visual media data according to the similarity between the visual media data.

[0156] In the above visual media data deduplication processing device, by separately extracting visual features from at least two visual media data to obtain the visual features of each visual media data, where the visual features include image features and text region features. By simultaneously considering the text region features on the visual media data, it is avoided that visual media data with the same image features and different text region features are classified as the same visual media data. By extracting text information from each visual media data to obtain the text content features of each visual media data, it is then possible to perform similarity analysis on at least two visual media data based on the visual features and the text content features, obtain the similarity between the visual media data, and perform deduplication processing on at least two visual media data according to the similarity between the visual media data. It realizes the combination consideration from multiple perspectives to improve the accuracy of the calculated similarity degree between videos, avoids the situation of missing duplicate videos without deduplication, and at the same time, by adopting the method of comprehensive consideration from multiple perspectives, it can also avoid the problem of over - attention to interference factors in the traditional deduplication method based on md5 values, improving the video deduplication rate and deduplication processing efficiency of the video platform.

[0157] In one embodiment, the visual feature extraction module is further configured to:

[0158] When it is detected that there is text in the visual media data, based on the trained feature extraction network, respectively perform text region position detection on at least two visual media data to generate the text region features of the visual media data; based on the trained feature extraction network, respectively perform image feature extraction on at least two visual media data to obtain the image features of at least two visual media data.

[0159] In one embodiment, the similarity analysis module is further configured to:

[0160] Based on the image features and text region features, perform visual similarity calculation to generate the visual similarity values of at least two visual media data; based on the text content features, perform text similarity calculation to generate the text similarity values of at least two visual media data; synthesize the visual similarity values and the text similarity values to obtain the similarity between the visual media data.

[0161] In one embodiment, a visual media data deduplication processing device is further provided, which further includes a video frame extraction module for performing video frame extraction on at least two video data to obtain video frames;

[0162] The similarity between the visual media data is the similarity between the video frames, and the deduplication processing module is further configured to:

[0163] Obtain a preset similarity threshold, and determine whether the similarity between video frames is greater than the preset similarity threshold; when it is determined that the similarity between video frames is greater than the preset similarity threshold, determine the current two video frames as a pair of similar video frames; obtain the number of pairs of similar video frames in the preset pairs of video frames extracted from any two video data; determine whether there are duplicate videos in the current two video data according to the preset threshold of pairs of similar video frames and the number of pairs of similar video frames; when it is determined that there are duplicate videos, perform video deduplication processing on the current two video data.

[0164] In one embodiment, there is provided a visual media data deduplication processing device, further comprising a feature extraction network training module for:

[0165] Randomly collect a set of visual media data samples; obtain different region images of the visual media data in the set of visual media data samples as training positive samples of visual media data; the region images in the training positive samples of visual media data include region images with text region features; obtain the same region images of different visual media data in the set of visual media data samples as training negative samples of visual media data; train the initial feature extraction network according to the training positive samples of visual media data and the training negative samples of visual media data to obtain a trained feature extraction network.

[0166] In one embodiment, the text content feature extraction module is further used for:

[0167] Perform image preprocessing on each visual media data to obtain a preprocessed image frame to be recognized; perform character segmentation and character recognition on the image frame to be recognized to obtain the segmented characters in sequence; perform dimensionality reduction processing and feature extraction on the segmented characters to obtain character features; perform feature classification and content recognition based on the character features to obtain the text content features of the visual media data.

[0168] In one embodiment, the video frame extraction module is further used for

[0169] Obtain target video information, and according to the target video information, perform key frame extraction on at least two video data to obtain video key frames; or perform timed video frame extraction on at least two video data at a preset frame extraction time interval to obtain corresponding fixed video frames.

[0170] For the specific limitations of the visual media data deduplication processing device, reference can be made to the limitations of the visual media data deduplication processing method in the above text, which will not be elaborated here. Each module in the above visual media data deduplication processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0171] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 8 . The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as visual media data, image features, text region features, text content features, and the similarity between visual media data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for deduplicating visual media data.

[0172] Those skilled in the art can understand that Figure 8 the structure shown in

[0173] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0174] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0175] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0176] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The above computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0177] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0178] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for deduplicating visual media data, characterized in that, The method includes: Performing visual feature extraction on at least two visual media data respectively to obtain the visual features of each of the visual media data, where the visual features include image features and text region features; Performing text information extraction on each of the visual media data to obtain the text content features of each of the visual media data; Based on the image features and the text region features, calculating visual similarity to generate a visual similarity value of the at least two visual media data, calculating text similarity based on the text content features to generate a text similarity value of the at least two visual media data, and synthesizing the visual similarity value and the text similarity value to obtain the similarity between the visual media data; Performing duplicate removal processing on the at least two visual media data according to the similarity between the visual media data.

2. The method according to claim 1, wherein The performing visual feature extraction on at least two visual media data respectively to obtain the visual features of each of the visual media data includes: When it is detected that there is text in the visual media data, based on the trained feature extraction network, performing text region position detection on at least two visual media data respectively to generate the text region features of the visual media data; Based on the trained feature extraction network, performing image feature extraction on at least two visual media data respectively to obtain the image features of the at least two visual media data.

3. The method according to claim 1 or 2, characterized in that, The visual media data includes video data; before performing visual feature extraction on at least two visual media data respectively to obtain the visual features of each of the visual media data, the method further includes: performing video frame extraction on at least two video data to obtain video frames; The similarity between the visual media data is the similarity between video frames; the performing duplicate removal processing on the at least two visual media data according to the similarity between the visual media data includes: Obtaining a preset similarity threshold and determining whether the similarity between the video frames is greater than the preset similarity threshold; When it is determined that the similarity between the video frames is greater than the preset similarity threshold, determining the current two video frames as a pair of similar video frames; Obtaining the number of pairs of similar video frames in the preset pair of video frames extracted from any two of the video data; Determining whether there are duplicate videos in the current two video data according to the preset threshold of pairs of similar video frames and the number of pairs of similar video frames; When it is determined that there are duplicate videos, performing video duplicate removal processing on the current two video data.

4. The method according to claim 2, wherein Obtaining the trained feature extraction network includes: Randomly collecting a sample set of visual media data; Obtaining different region images of the visual media data in the sample set of visual media data as positive samples for training visual media data; the region images in the positive samples for training visual media data include region images with text region features; Obtaining the same region images of different visual media data in the sample set of visual media data as negative samples for training visual media data; Training the initial feature extraction network according to the positive samples for training visual media data and the negative samples for training visual media data to obtain the trained feature extraction network.

5. The method according to claim 1 or 2, characterized in that, Performing text information extraction on each of the visual media data to obtain the text content features of each of the visual media data includes: Performing image preprocessing on each of the visual media data to obtain a preprocessed image frame to be recognized; Based on the image frame to be recognized, performing character segmentation and character recognition to sequentially obtain the segmented characters; Performing dimensionality reduction processing and feature extraction on the segmented characters to obtain character features; Based on the character features, performing feature classification and content recognition to obtain the text content features of the visual media data.

6. The method according to claim 3, wherein The video frame extraction for at least two video data to obtain video frames includes: Obtaining target video information, and based on the target video information, performing key frame extraction on the at least two video data to obtain video key frames; or Performing timed video frame extraction on the at least two video data at a preset frame extraction time interval to obtain corresponding fixed video frames.

7. A visual media data deduplication processing device, characterized in that, The device includes: A visual feature extraction module, configured to perform visual feature extraction on at least two visual media data respectively to obtain the visual features of each of the visual media data, where the visual features include image features and text region features; A text content feature extraction module, configured to perform text information extraction on each of the visual media data to obtain the text content features of each of the visual media data; A similarity analysis module, configured to perform visual similarity calculation based on the image features and the text region features to generate a visual similarity value of the at least two visual media data, perform text similarity calculation based on the text content features to generate a text similarity value of the at least two visual media data, and combine the visual similarity value and the text similarity value to obtain the similarity between the visual media data; A duplicate removal processing module, configured to perform duplicate removal processing on the at least two visual media data according to the similarity between the visual media data.

8. The device according to claim 7, characterized in that The visual feature extraction module is further configured to: When it is detected that there is text in the visual media data, based on a trained feature extraction network, perform text region position detection on at least two visual media data respectively to generate the text region features of the visual media data; Based on the trained feature extraction network, perform image feature extraction on at least two visual media data respectively to obtain the image features of the at least two visual media data.

9. The device according to claim 7 or 8, characterized in that, The visual media data includes video data; the device further includes a video frame extraction module, configured to perform video frame extraction on at least two video data to obtain video frames; The similarity between the visual media data is the similarity between video frames; the deduplication processing module is further configured to: obtain a preset similarity threshold, and determine whether the similarity between the video frames is greater than the preset similarity threshold; when it is determined that the similarity between the video frames is greater than the preset similarity threshold, determine the current two video frames as a pair of similar video frames; obtain the number of pairs of similar video frames in the preset pair of video frames extracted from any two of the video data; determine whether there are duplicate videos in the current two video data according to the preset threshold of pairs of similar video frames and the number of pairs of similar video frames; when it is determined that there are duplicate videos, perform video deduplication processing on the current two video data.

10. The device according to claim 8, characterized in that, The device further includes a feature extraction network training module, configured to: Randomly collect a sample set of visual media data; obtain different regional images of the visual media data in the sample set of visual media data as positive samples for training visual media data; the regional images in the positive samples of visual media data include regional images with text area features; obtain the same regional images of different visual media data in the sample set of visual media data as negative samples for training visual media data; train an initial feature extraction network according to the positive samples for training visual media data and the negative samples for training visual media data to obtain a trained feature extraction network.

11. The device according to claim 7 or 8, characterized in that, The text content feature extraction module is further configured to: Perform image preprocessing on each of the visual media data to obtain a preprocessed image frame to be recognized; perform character segmentation and character recognition on the image frame to be recognized to obtain the segmented characters in sequence; Perform dimensionality reduction processing and feature extraction on the segmented characters to obtain character features; Perform feature classification and content recognition based on the character features to obtain the text content features of the visual media data.

12. The device according to claim 9, wherein, The video frame extraction module is further configured to: Obtain target video information, and according to the target video information, extract key frames from the at least two video data to obtain video key frames; or perform timed video frame extraction on the at least two video data at a preset frame extraction time interval to obtain corresponding fixed video frames.

13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method and device for determining similar data, electronic equipment and storage medium

    CN110414625A

  • Short video detection and multi-classification method and device and storage medium

    CN113779308A