Video data processing method and apparatus, storage medium, and electronic device

By clustering video data and obtaining multi-dimensional similarity to assist manual annotation, the problems of low efficiency and low accuracy in video annotation are solved, and efficient and accurate video annotation is achieved.

CN114329059BActive Publication Date: 2025-11-21TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111465972.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-11-21
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Existing technologies for video annotation are inefficient and inaccurate. Furthermore, manual annotation relies on human resources and cannot meet the demands of increasing video data volume and complexity. The single clustering method prevents annotators from effectively utilizing clustering information, resulting in time-consuming and inefficient annotation.

Method used

By clustering video data, the similarity of video pairs in multiple dimensions is obtained. The clustering results are used to assist manual annotation, including the extraction of labels and similarity analysis of image, audio and text dimensions, to form clustering results for batch annotation.

Benefits of technology

It improves the efficiency and accuracy of manual annotation, reduces annotation pressure, saves labor costs, and achieves efficient annotation of video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114329059B_ABST
    Figure CN114329059B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a video data processing method and device, a storage medium and an electronic device, which are applied to the field of artificial intelligence. The method comprises obtaining at least two videos. For each video in the at least two videos, the labels of at least two dimensions in the video are determined. For a video pair formed by the at least two videos, the video pair is clustered based on the dimension similarity of the at least two dimensions corresponding to the video pair, to obtain a clustering result corresponding to the at least two videos, and the dimension similarity is used to represent the coincidence degree of the labels of the same dimension in the two videos of the video pair. The present application can cluster video data, so that clustering information can be obtained during manual annotation, so that manual annotation can be batched according to the clustering result, or annotation can be performed by referring to the clustering information, thereby solving the problems of low efficiency, low accuracy and low production capacity of current manual annotation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of artificial intelligence, and in particular, to a video data processing method and device, a storage medium, and an electronic device. BACKGROUND

[0002] With the rapid development of the Internet and multimedia technology, vivid digital videos have gradually replaced monotonous text information and become one of the important ways of information dissemination in people's online life. In the face of a large number of videos on the Internet, it has become a problem that people pay more and more attention to whether the required video can be found in a short time. Video search, video recommendation, video analysis and other applications that rely on video annotation have also gradually received more and more attention.

[0003] However, with the development of these applications, the data volume of videos increases rapidly, the difficulty of video annotation increases, and the annotation speed and accuracy are difficult to meet the development needs of related applications. SUMMARY

[0004] In order to assist in improving the efficiency and accuracy of annotation, reducing the pressure of annotation, and saving the cost of annotation, embodiments of the present application provide a video data processing method and device, a storage medium, and an electronic device.

[0005] In one aspect, a video data processing method is provided, and the method comprises:

[0006] obtaining at least two videos;

[0007] determining, for each video of the at least two videos, labels of at least two dimensions in the video;

[0008] obtaining, for a video pair formed by the at least two videos, a dimension similarity corresponding to the at least two dimensions, and performing clustering based on the dimension similarity corresponding to the at least two dimensions, to obtain a clustering result corresponding to the at least two videos, the dimension similarity being used to represent a coincidence degree of labels in the same dimension of the two videos of the video pair.

[0009] In another aspect, a video data processing device is provided, and the device comprises:

[0010] a video obtaining module configured to obtain at least two videos;

[0011] a dimension label obtaining module configured to determine, for each video of the at least two videos, labels of at least two dimensions in the video;

[0012] The clustering module is configured to obtain dimension similarity of the video pair in the at least two dimensions, and perform clustering based on the dimension similarity of the video pair in the at least two dimensions to obtain a clustering result corresponding to the at least two videos, wherein the dimension similarity is used to represent coincidence of labels of the same dimension in the two videos of the video pair.

[0013] In another aspect, an embodiment of the present application provides a computer readable storage medium, which stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the video data processing method.

[0014] In another aspect, an embodiment of the present application provides an electronic device, which comprises at least one processor and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the video data processing method by executing the instructions stored in the memory.

[0015] In another aspect, an embodiment of the present application provides a computer program product, which comprises a computer program or instructions, and the computer program or instructions are executed by a processor to implement the video data processing method.

[0016] The video data processing method provided by the present application can obtain clustering information during manual labeling, so that manual labeling can be performed in batches according to the clustering result, or manual labeling can be performed by referring to the clustering information, thereby solving the problems of low efficiency, low accuracy and low productivity of manual labeling. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the related art, the drawings needed in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0018] Figure 1 is a schematic diagram of labeling provided in the related art by an embodiment of the present specification;

[0019] Figure 2 is a schematic diagram of a feasible implementation framework of the video data processing method provided by an embodiment of the present specification;

[0020] Figure 3 is a flowchart of a video data processing method provided by an embodiment of the present application;

[0021] Figure 4 is a schematic diagram of importing a video provided by an embodiment of the present application;

[0022] Figure 5 is a schematic diagram of audio recognition provided by an embodiment of the present application;

[0023] Figure 6 is a schematic diagram of character recognition provided by an embodiment of the present application;

[0024] Figure 7 is a schematic diagram of a label extraction result corresponding to a picture dimension provided by an embodiment of the present application;

[0025] Figure 8 is a schematic diagram of a video label provided by an embodiment of the present application;

[0026] Figure 9 is an interactive flowchart of a video data processing method provided by an embodiment of the present application;

[0027] Figure 10 is a logic flowchart of a video data processing method provided by an embodiment of the present application;

[0028] Figure 11 is a block diagram of a video data processing apparatus provided by an embodiment of the present application;

[0029] Figure 12 is a hardware structure schematic diagram of a device for implementing the method provided by an embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0031] It should be noted that the terms "first", "second", etc. in the description of the embodiments of the present application and the claims and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application are further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, and do not limit the embodiments of the present application.

[0033] Hereinafter, the terms "first" and "second" are used only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features. In the description of the embodiments, unless otherwise specified, the meaning of "a plurality of" is two or more. In order to facilitate understanding of the technical solutions and the technical effects generated by the above-mentioned technical solutions of the embodiments of the present application, the embodiments of the present application first explain the related professional terms:

[0034] Artificial intelligence (Artificial Intelligence, AI) is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0035] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software level technology. Artificial intelligence basic technology generally includes, such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, automatic driving, intelligent transportation and other major directions.

[0036] Machine learning (ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It is a specialized study of how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. Its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and example-based learning.

[0037] Convolutional Neural Networks (CNN) is a class of feedforward neural networks that contain convolutional computation and have a deep structure, and is one of the representative algorithms of deep learning. Convolutional neural networks have the ability to represent learning, and can classify input information according to their hierarchical structure, so they are also called translation-invariant artificial neural networks.

[0038] Video OCR recognition: Video OCR recognition mainly includes three links of front-end video information collection and transmission, middle video detection and back-end analysis processing. Video recognition needs a clear and stable video signal provided by the front-end video capture camera, and the quality of the video signal will directly affect the effect of video recognition. Then through the embedded intelligent analysis module, using OCR (Optical Character Recognition) technology, the video picture is recognized, detected, analyzed and filtered out of interference, and the abnormal situation in the video picture is marked with target and trajectory. The intelligent video analysis module is an algorithm based on artificial intelligence and pattern recognition principles.

[0039] Equal-interval frame extraction: Equal-interval frame extraction is a common way of extracting frames from a video. It is a process of simulating taking a photo every certain time and splicing it to form a video (i.e. low-speed shooting) by extracting a certain number of frames at intervals in a video.

[0040] Image clustering: the process of image clustering is essentially a knowledge-based image understanding process, and also the continuation and development of human visual discrimination of images. The research on image clustering based on visual features is an important way to solve visual image problems, and is also an interdisciplinary research method that gathers computer vision, image processing, data mining and other research fields. For specific problems and users, a variety of representative clustering algorithms have been proposed and widely used in pattern recognition, biological information, image processing and data mining fields.

[0041] In the related art, artificial intelligence is relied on to provide users with video search, video analysis, video recommendation, video analysis and other video content services. The quality of these video content services depends largely on the quality of the corresponding artificial intelligence model, and the quality of the artificial intelligence model depends on video annotation. With the rapid development of artificial intelligence, there is a large gap in the demand for annotated corpus, and a large amount of specified video annotation data needs to be trained to make the technology more mature.

[0042] With the increase of video data volume, the increase of video content complexity, the improvement of artificial intelligence model quality requirements and the improvement of artificial intelligence model complexity, higher requirements are put forward for video annotation. Traditional manual annotation relies on manpower to annotate videos, and neither the annotation efficiency nor the annotation accuracy can meet the requirements. The accuracy of machine annotation is limited and it is difficult to be used independently without manual annotation. At present, manual annotation is still indispensable. In the related art, video clustering can be used to assist manual annotation, but the clustering method in the related art is relatively single and cannot be directly clustered by the video itself, resulting in a certain loss of clustering effect. When the annotator annotates, he has no perception of clustering and it is difficult to use clustering information to assist annotation. In video annotation, unprocessed video information is very scattered and information is complex. The annotator needs to constantly change his thinking when annotating, which is very time-consuming. Please refer to Figure 1 which shows a schematic diagram of annotation in the related art. The annotation content of the image needs to completely rely on manual analysis of the picture, and the annotator lacks reference knowledge, resulting in low annotation efficiency, poor quality and low productivity.

[0043] In order to assist in improving the efficiency and accuracy of manual annotation, reducing the pressure of manual annotation and saving the cost of manual annotation, an embodiment of the present application proposes a video data processing method, which clusters video data so that clustering information can be obtained when manual annotation is performed, so that manual annotation can be performed in batches according to the clustering result, or the clustering information can be referred to for annotation.

[0044] The method provided by the embodiments of the present application can be related to a blockchain, that is, the method provided by the embodiments of the present application can be implemented based on a blockchain, or the data involved in the method provided by the embodiments of the present application can be stored based on a blockchain, or the execution subject of the method provided by the embodiments of the present application can be located in a blockchain. The blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. The blockchain (Blockchain) is essentially a decentralized database, and is a series of data blocks associated using a cryptographic method, each data block containing information of a batch of network transactions, for verifying the validity (anti-fake) of the information and generating the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.

[0045] The blockchain underlying platform can include user management, basic services, smart contracts, and operation monitoring processing modules. The user management module is responsible for identity information management of all blockchain participants, including maintaining public and private key generation (account management), key management, and user real identity and blockchain address correspondence maintenance (permission management), and in the case of authorization, supervising and auditing the transaction of certain real identities, providing risk control rule configuration (risk audit); the basic service module is deployed on all blockchain node devices to verify the validity of business requests, and after consensus, the valid requests are recorded on the storage. For a new business request, the basic service first performs interface adaptation analysis and authentication processing (interface adaptation), then encrypts the business information through a consensus algorithm (consensus management), and after encryption, the complete and consistent information is transmitted to the shared ledger (network communication) and recorded and stored; the smart contract module is responsible for contract registration and issuance, contract triggering and contract execution. Developers can define contract logic through a certain programming language, publish it to the blockchain (contract registration), call keys or other events to trigger execution according to the logic of the contract terms, complete the contract logic, and also provide contract upgrade and cancellation functions; the operation monitoring module is mainly responsible for deployment, configuration modification, contract setting, cloud adaptation during product release, and real-time state visualization output during product operation, such as alarm, monitoring network conditions, and monitoring node device health status.

[0046] The platform product service layer provides basic capabilities and implementation frameworks for typical applications. Developers can stack business characteristics based on these basic capabilities to complete blockchain implementation of business logic. The application service layer provides application services based on the blockchain scheme for business participants to use.

[0047] Please refer to Figure 2 , Figure 2 is a feasible implementation framework schematic diagram of the video data processing method provided by the embodiments of the present application, as shown in Figure 1As shown, the implementation framework can at least include terminal device 01, data processing server 02. Among them, the terminal device 01 can be a device located in the Internet, which can display the clustering result for the user by interacting with the data server 02, and allow the user to label the video based on the clustering result. The terminal device 01 includes but is not limited to mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle-mounted terminals, etc.

[0048] The data processing server 02 can obtain at least two videos. For each of the at least two videos, determine the labels of at least two dimensions in the video. For the video pair formed by the at least two videos, based on the dimension similarity of the at least two dimensions corresponding to the video pair, clustering is performed to obtain the clustering result corresponding to the at least two videos, and the dimension similarity is used to represent the coincidence degree of the labels of the same dimension in the two videos of the video pair.

[0049] The following describes a video data processing method of an embodiment of the application, Figure 3 The flowchart of the video data processing method provided by the embodiment of the application is shown, and the method operation steps described in the embodiment or the flowchart can include more or fewer operation steps based on conventional or non-creative labor. The order of the steps listed in the embodiment is only one of the many execution orders, and does not represent the only execution order. When the system, terminal device or server product in practice is executed, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method order shown in the embodiment or the drawing. The above method can include:

[0050] S101. Obtain at least two videos.

[0051] In the embodiment of the application, the address of the at least two videos can be obtained to access the at least two videos, or the at least two videos can be obtained directly. The application does not limit the source of the at least two videos, which can be imported by the user or captured in the cloud or the Internet. Please refer to Figure 4 which shows a schematic diagram of importing videos. The user can import videos in the form of a video address list.

[0052] S102. For each of the at least two videos, determine the labels of at least two dimensions in the video.

[0053] For each of the at least two videos, at least two of the labels corresponding to the picture dimension, the labels corresponding to the audio dimension and the labels corresponding to the text dimension in the video are determined.

[0054] Specifically, when the labels of at least two dimensions include the labels corresponding to the image dimension, the above determination of the labels of at least two dimensions in the video includes: extracting at least two video frames in the video; performing image recognition on each video frame to obtain the labels corresponding to the image dimension.

[0055] When the labels in at least two dimensions include the label corresponding to the audio dimension, determining the labels in at least two dimensions of the video includes: classifying and recognizing the audio information in the video to obtain the label corresponding to the audio dimension. Specifically, audio features within video frames can be extracted from the imported video file, the audio features of each video frame can be classified and recognized, and the confidence score of each classification result can be scored. The classification corresponding to the highest score is used as the audio feature label of the video frame, and the audio feature labels corresponding to each video frame are deduplicated to obtain the label corresponding to the audio dimension. If the output value is empty, it means that the video has no audio information. This application does not limit the method of audio feature extraction. For example, it can be extracted based on a convolutional neural network, an encoder network, or frequency domain sampling. When classifying and recognizing based on audio features, classification prediction can be performed based on a decoder network or a classification network. This application does not limit the specific execution method of classification and recognition. The obtained classification result can include a classification label and a classification confidence score. The classification confidence score represents the confidence score of the classification result and the probability that the audio belongs to the category pointed to by the classification label.

[0056] Of course, the embodiments of this application can classify audio at least by one level; please refer to [the relevant documentation]. Figure 5 It shows a schematic diagram of audio recognition, for Figure 5 The video file can be classified into two levels. The first level of classification is used to identify whether the audio in the video is speech or a song. The second level of classification is used to identify which language the audio in the video is, whether it is a solo voice or a broadcast voice, if it is a song, and whether it is instrumental music.

[0057] When the labels in at least two of the aforementioned dimensions include labels corresponding to the text dimension, determining the labels in at least two dimensions of the aforementioned video includes: extracting text information from the aforementioned video to obtain the labels corresponding to the aforementioned text dimension. Specifically, the imported video file can be processed using video OCR technology to extract all text keyframes within the video, and the text information corresponding to each keyframe can be identified using OCR technology. Based on all the text information, the corresponding text information for the video is generated. If the output value is empty, it means that the video has no text information output. Please refer to [reference needed]. Figure 6 It shows a diagram of text recognition, which identifies each text keyframe and combines the recognition results to obtain the label corresponding to the text dimension.

[0058] In the case that the at least two dimension labels include a picture dimension label, the picture recognition of each video frame to obtain the picture dimension label includes:

[0059] S201. Picture recognition is performed on each video frame to obtain at least one recognition result corresponding to the video frame, the recognition result including a picture label and a confidence degree corresponding to the picture label.

[0060] The embodiments of the present application do not limit the way of video frame extraction, for example, the video frames in step S201 can be extracted according to equal interval frame extraction or a preset key frame extraction algorithm. The embodiments of the present application do not limit the key frame extraction method, for example, the pixel frame average method or the histogram frame average method can be used to extract the key frame. For each video frame, picture recognition is performed to obtain at least one recognition result corresponding to the video frame, the recognition result including a picture label and a confidence degree corresponding to the picture label. The embodiments of the present application do not limit the specific method of picture recognition, for example, artificial intelligence technology can be used to extract feature information in the picture, and then classification is performed based on the feature information to obtain the recognition result.

[0061] S202. The picture label with the confidence degree meeting the first requirement is determined as the picture label corresponding to the video frame.

[0062] The embodiments of the present application do not limit the first requirement, for example, each picture label can be arranged in descending order of confidence degree, and the first preset number of picture labels are determined as the picture labels corresponding to the video frame. Of course, the present application does not limit the first preset number, which can be set according to actual conditions, for example, the first preset number can be set to 5.

[0063] S203. The picture labels corresponding to each video frame are counted, and the picture label with the repetition number meeting the second requirement is determined as the label corresponding to the picture dimension.

[0064] The embodiments of the present application do not limit the second requirement, for example, the picture label with the repetition number greater than the second preset number in the set formed by the picture labels corresponding to each video frame is determined as the label corresponding to the picture dimension. Of course, the present application does not limit the second preset number, which can be set according to actual conditions.

[0065] In another embodiment, if the total number of repeated times of the picture tags that meet the second requirement does not meet the third preset number of tags corresponding to the picture dimension, the picture tags that do not meet the second requirement can also be jointly sorted in descending order based on the repeated times and the maximum confidence, and the picture tags are selected according to the sorting result until the third preset number of tags corresponding to the picture dimension is met. In the joint descending sorting, the sorting priority of the repeated times is higher than that of the maximum confidence, and the maximum confidence refers to the maximum value of the confidence corresponding to the picture tag. Of course, the third preset number is not limited in the present application, and can be set according to the actual situation, for example, it can be set to 5.

[0066] In another embodiment, the picture tags corresponding to each video frame can be directly jointly sorted in descending order based on the repeated times and the maximum confidence, and the picture tags are selected according to the sorting result until the third preset number of tags corresponding to the picture dimension is met.

[0067] Please refer to Figure 7 which shows a schematic diagram of the picture tag extraction result corresponding to the picture dimension. Figure 7 In the picture tag extraction operation of the three video frames, the picture tags corresponding to the video frame 1 are "person", "text", and "blue", the picture tags corresponding to the video frame 2 are "person", "text", "blue", "red", and "balloon", and the picture tags corresponding to the video frame 3 are "person", "text", "blue", "red", and "motor". If the second preset number is 2, the repeated times of "person", "text", and "blue" are all 3, which meet the second requirement, and are selected as the tags corresponding to the picture dimension.

[0068] For the tags that do not meet the second requirement, "red", "motor", and "balloon", "red" appears twice and is ranked first. If the maximum confidence of "motor" is higher than that of "balloon", "motor" is ranked second and "balloon" is ranked third. If the third preset number is 4, "person", "text", "blue", and "red" are selected as the tags corresponding to the picture dimension.

[0069] S103. For the video pair formed by the at least two videos, the dimension similarity corresponding to the at least two dimensions of the video pair is obtained, and clustering is performed based on the dimension similarity corresponding to the at least two dimensions of the video pair to obtain the clustering result corresponding to the at least two videos. The dimension similarity is used to represent the coincidence degree of the tags in the same dimension in the two videos of the video pair.

[0070] Specifically, the first video and the second video can be clustered based on the dimension similarity to obtain a clustering tendency corresponding to each video pair, the video pair including the first video and the second video, and the clustering tendency being used to indicate a probability that the first video and the second video belong to similar videos. The clustering result is obtained according to the clustering tendency corresponding to each video pair. For example, the at least two videos are clustered in descending order of the clustering tendency to obtain the clustering result. For example, for a video pair A formed by video 1 and video 2, the video 1 and the video 2 can be compared in the audio, picture and text dimensions to obtain the similarity of the video 1 and the video 2. If only one of the audio, picture and text dimensions is similar, the clustering tendency corresponding to the video pair A is determined to be s1; if two of the audio, picture and text dimensions are similar, the clustering tendency corresponding to the video pair A is determined to be s2; if all of the audio, picture and text dimensions are similar, the clustering tendency corresponding to the video pair A is determined to be s3, and obviously s3 is greater than s2 which is greater than s1. Taking video pairs A, B and C as examples, if the result of descending order based on the clustering tendency is B, A and C, the two videos in the video pair B are clustered first, then the two videos in the video pair A are clustered, and finally the two videos in the video pair C are clustered.

[0071] In one embodiment, the clustering of the first video and the second video based on the dimension similarity to obtain the clustering tendency corresponding to the video pair includes: determining a coincidence of the label of each dimension of the first video and the label of each dimension of the second video to obtain a coincidence index corresponding to each dimension. The clustering tendency corresponding to the video pair is obtained according to the coincidence index corresponding to each dimension.

[0072] In one embodiment, the determination of the coincidence of the label of each dimension of the first video and the label of each dimension of the second video to obtain the coincidence index corresponding to each dimension includes: obtaining a coincidence degree of the label of each dimension of the first video and the label of each dimension of the second video; and in response to a case that the coincidence degree is greater than a preset threshold corresponding to each dimension, determining the coincidence index corresponding to each dimension as a first index. Correspondingly, the obtaining of the clustering tendency corresponding to the video pair according to the coincidence index corresponding to each dimension includes: obtaining the clustering tendency corresponding to the video pair according to the number of the first index.

[0073] The clustering process is described in detail by taking the first video and the second video in the video pair as an example, which both include picture dimension corresponding tags, audio dimension corresponding tags and text dimension corresponding tags. Of course, if there is no audio information in the first video or the second video, the audio dimension corresponding tag can be empty, and if there is no text information in the first video or the second video, the text dimension corresponding tag can be empty.

[0074] Taking the picture dimension as an example, the number of coincidences of the picture dimension corresponding tags of the first video and the picture dimension tags of the second video is N1, and the preset threshold value of the picture dimension is M1. If N1 is greater than M1, it can be considered that the first video and the second video are similar in the picture dimension, and the coincidence index corresponding to the picture dimension is determined as the first index Y.

[0075] Taking the audio dimension as an example, the number of coincidences of the audio dimension corresponding tags of the first video and the audio dimension tags of the second video is N2, and the preset threshold value of the audio dimension is M2. If N2 is greater than M2, it can be considered that the first video and the second video are similar in the audio dimension, and the coincidence index corresponding to the audio dimension is determined as the first index Y.

[0076] Taking the text dimension as an example, the number of coincidences of the text dimension corresponding tags of the first video and the text dimension tags of the second video is N3, and the preset threshold value of the text dimension is M3. If N3 is greater than M3, it can be considered that the first video and the second video are similar in the text dimension, and the coincidence index corresponding to the text dimension is determined as the first index Y.

[0077] From the above, it can be seen that for the video pair formed by the first video and the second video, the number of the first index Y is most likely to be 3 and least likely to be 0. For each video pair, the number of the first index Y possessed by it can be counted. Only when the number of the first index Y is greater than or equal to the preset clustering threshold value, the clustering can be performed, otherwise the clustering cannot be performed. The video pairs with more first index Y can be preferentially clustered into clusters. Any video and video, video and cluster, or cluster and cluster can continue to be clustered. The specific clustering method is to select one video from the two parties to be clustered to form a video pair, and determine the number of the first index of the video pair according to the above content. If the number of the first index of the video pair is greater than or equal to the above clustering threshold value, the two parties can be clustered.

[0078] Taking the picture dimension as an example, Figure 8For example, there are videos A, B, C, forming video pairs AB, AC, BC, preset thresholds M1, M2, M3 can be set as 4, 3, 5, for video pair AB, the picture dimension under the coincidence number N1 is 5, therefore, the coincidence index corresponding to the picture dimension is the first index Y, the audio dimension under the coincidence number N2 is 4, therefore, the coincidence index corresponding to the audio dimension is also the first index Y, the coincidence number N3 under the text dimension is 2, the coincidence index corresponding to the text dimension is not Y, therefore, the number of the first index Y possessed by the video pair AB is 2. Similarly, the number of the first index Y possessed by the video pairs AC and BC is 0 and 0 respectively. If the aggregation threshold is 1, only the video pair AB is clustered.

[0079] Of course, after obtaining the clustering result, the clustering result can also be displayed to the user, and the user can also be provided with clustering information for reference when labeling the video. Please refer to Figure 8 , which shows a clustering result display diagram. By displaying the clustering result and the clustering information, the user can be prompted and referenced when labeling the video.

[0080] In a specific embodiment, as Figure 9 shown, the video data processing method in the embodiment of the application can specifically include the following steps:

[0081] Step one: the user imports the video data to be processed into the data processing server 02 through the terminal device 01, and the data processing server 02 loads the video data. Of course, in some embodiments, the video content can also be displayed.

[0082] Step two: equal-interval frame extraction. The data processing server 02 performs equal-interval frame extraction on the video, specifically, frame extraction can be performed every 0.1 second, and a plurality of pictures are obtained.

[0083] Step three: picture dimension corresponding label extraction. The data processing server 02 extracts the label corresponding to each picture, specifically, only the top 5 labels with the highest confidence can be extracted, and after the labels of each picture are summarized, a plurality of labels with the most repeated times are extracted.

[0084] Step four: audio dimension corresponding label extraction. The audio information of the video is extracted, and the audio label is extracted. Specifically, the audio feature corresponding to the video can be obtained by extracting the feature according to the voiceprint and audio content, and the audio label prediction is performed, and of course, the null value can be displayed under the condition of no audio information.

[0085] Step five: text dimension corresponding label extraction. The text information of the video is extracted to obtain the label corresponding to the text dimension. Specifically, the text key frame can be extracted and the text content can be recognized according to the OCR, and the text content of each key frame is de-duplicated and summarized to output the label corresponding to the text dimension.

[0086] The embodiments of the present application do not limit the execution order of steps three, four and five, which can be executed sequentially, in parallel, or in any order and then executed according to the order.

[0087] Step six: clustering based on the extracted labels. Specifically, each dimension of a video can be matched with the same dimension of other videos, and if 80% of the labels are the same, it is determined that the two videos match in this dimension. If there is one or more than one dimension matching in the three dimension features of different videos, it is determined that they are similar videos and are clustered.

[0088] Step seven: display the clustering results. Specifically, videos that have been successfully clustered can be displayed first, and then videos that have not been clustered. Specifically, clustering can be performed according to the dimension matching order, with three dimensions being greater than two dimensions and one dimension. If a cluster contains two or more videos, it is clustered. Videos that have been successfully clustered are displayed first. If a cluster contains only one video, it is a failed cluster. Videos that have failed to cluster can be displayed after the successfully clustered videos.

[0089] Step eight: manually label the clustering results and output the results.

[0090] Of course, steps one to eight above can be automatically executed on the data processing server 02, or can be executed by relying on the interaction between the terminal device 01 and the data processing server 02.

[0091] Please refer to Figure 10 , which shows a logic flowchart corresponding to Figure 9 Step. After the video is entered, the feature corresponding to the video picture information is extracted by the equidistant frame extraction technology, and the label pointed by the feature corresponding to the picture information is extracted. Then, the label pointed by the audio is obtained by the audio recognition technology. Further, the text information in the video is extracted by the video OCR technology, and the label pointed by the text is obtained. Based on the three labels, a batch of videos can be clustered and the clustering results can be output. The video data processing method proposed in the embodiments of the present application clusters the video data, so that the clustering information can be obtained during manual labeling, so that the manual labeling can be batched according to the clustering results, or the labeling can be performed according to the clustering information, thereby solving the problems of low efficiency, low accuracy and low capacity of manual labeling.

[0092] Please refer to Figure 11 , which shows a block diagram of a video data processing device in the embodiments of the present application. The above device comprises:

[0093] The video acquisition module 101 is configured to acquire at least two videos.

[0094] The dimension label obtaining module 102 is configured to determine labels of at least two dimensions in each of the at least two videos.

[0095] The clustering module 103 is configured to obtain dimension similarities of the at least two dimensions corresponding to the video pair, and perform clustering based on the dimension similarities of the at least two dimensions corresponding to the video pair, to obtain a clustering result corresponding to the at least two videos. The dimension similarity is used to represent coincidence of labels of the same dimension in the two videos of the video pair.

[0096] In an embodiment, the dimension label obtaining module is configured to perform the following operations:

[0097] The dimension label obtaining module is configured to determine at least two of the labels of the picture dimension, the labels of the audio dimension, and the labels of the text dimension in each of the at least two videos.

[0098] In an embodiment, when the labels of the at least two dimensions include the labels of the picture dimension, the dimension label obtaining module is configured to extract at least two video frames in the video; perform picture recognition on each video frame to obtain the labels of the picture dimension.

[0099] When the labels of the at least two dimensions include the labels of the audio dimension, the dimension label obtaining module is configured to perform classification recognition on audio information in the video to obtain the labels of the audio dimension.

[0100] When the labels of the at least two dimensions include the labels of the text dimension, the dimension label obtaining module is configured to extract text information in the video to obtain the labels of the text dimension.

[0101] In an embodiment, when the labels of the at least two dimensions include the labels of the picture dimension, the dimension label obtaining module is configured to perform the following operations: perform picture recognition on each video frame to obtain at least one recognition result corresponding to the video frame, wherein the recognition result includes a picture label and a confidence degree corresponding to the picture label.

[0102] The picture label with the confidence degree meeting a first requirement is determined as the picture label corresponding to the video frame.

[0103] The picture labels corresponding to each video frame are counted, and a picture label with a repeated number meeting a second requirement is determined as the label of the picture dimension.

[0104] In an embodiment, the clustering module is configured to perform the following operations:

[0105] For each video pair determined according to the at least two videos, the first video and the second video are clustered based on dimension similarity to obtain a clustering tendency corresponding to the video pair, the video pair including the first video and the second video, and the clustering tendency is used to indicate a probability that the first video and the second video belong to similar videos.

[0106] The clustering result is obtained according to the clustering tendency corresponding to each video pair.

[0107] In an embodiment, the clustering module is configured to perform the following operations:

[0108] The coincidence of the label of the first video in each dimension and the label of the second video in the each dimension is determined to obtain a coincidence indicator corresponding to the each dimension;

[0109] The clustering tendency corresponding to the video pair is obtained according to the coincidence indicator corresponding to each dimension.

[0110] In an embodiment, the clustering module is configured to perform the following operations:

[0111] The coincidence degree of the label of the first video in each dimension and the label of the second video in the each dimension is obtained, and in response to the coincidence degree being greater than a preset threshold corresponding to the each dimension, the coincidence indicator corresponding to the each dimension is determined as a first indicator;

[0112] The clustering tendency corresponding to the video pair is obtained according to the coincidence indicator corresponding to each dimension, including obtaining the clustering tendency corresponding to the video pair according to the number of the first indicators.

[0113] In an embodiment, the clustering module is configured to perform the following operations:

[0114] The at least two videos are clustered in descending order of the clustering tendency to obtain the clustering result.

[0115] Embodiments of the present application also provide a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the above-mentioned video data processing method.

[0116] Embodiments of the present application also provide a computer readable storage medium, which can store a plurality of instructions. The instructions can be adapted to be loaded and executed by a processor to perform the above-mentioned video data processing method.

[0117] In an embodiment, the above method for processing video data comprises:

[0118] obtaining at least two videos;

[0119] for each of the at least two videos, determining labels of at least two dimensions in the video;

[0120] for a video pair formed by the at least two videos, obtaining dimension similarities of the video pair in the at least two dimensions, and performing clustering based on the dimension similarities of the video pair in the at least two dimensions to obtain a clustering result corresponding to the at least two videos, wherein the dimension similarity is used to represent a coincidence degree of labels in the same dimension in the two videos of the video pair.

[0121] In an embodiment, the above determining, for each of the at least two videos, labels of at least two dimensions in the video comprises:

[0122] for each of the at least two videos, determining at least two of a label corresponding to a picture dimension, a label corresponding to an audio dimension, and a label corresponding to a text dimension in the video.

[0123] In an embodiment, when the labels of the at least two dimensions include the label corresponding to the picture dimension, the above determining labels of at least two dimensions in the video comprises: extracting at least two video frames in the video; and performing picture recognition on each video frame to obtain the label corresponding to the picture dimension.

[0124] when the labels of the at least two dimensions include the label corresponding to the audio dimension, the above determining labels of at least two dimensions in the video comprises: performing classification recognition on audio information in the video to obtain the label corresponding to the audio dimension.

[0125] when the labels of the at least two dimensions include the label corresponding to the text dimension, the above determining labels of at least two dimensions in the video comprises: extracting text information in the video to obtain the label corresponding to the text dimension.

[0126] In an embodiment, when the labels of the at least two dimensions include the label corresponding to the picture dimension, the above performing picture recognition on each video frame to obtain the label corresponding to the picture dimension comprises:

[0127] performing picture recognition on each video frame to obtain at least one recognition result corresponding to the video frame, wherein the recognition result includes a picture label and a confidence degree corresponding to the picture label.

[0128] determining the picture label of the picture frame as the picture label corresponding to the video frame according to the first requirement;

[0129] counting the picture labels corresponding to the video frames, and determining the label corresponding to the picture dimension according to the picture label with the second requirement in the number of repetitions.

[0130] In one embodiment, the video pairs formed according to the at least two videos are clustered based on the dimension similarity in the at least two dimensions to obtain the clustering result corresponding to the at least two videos, including:

[0131] For each video pair determined according to the at least two videos, the first video and the second video are clustered based on the dimension similarity to obtain the clustering tendency degree corresponding to the video pair, the video pair includes the first video and the second video, and the clustering tendency degree is used to indicate the probability that the first video and the second video belong to similar videos;

[0132] The clustering result is obtained according to the clustering tendency degree corresponding to each video pair.

[0133] In one embodiment, the clustering of the first video and the second video based on the dimension similarity to obtain the clustering tendency degree corresponding to the video pair includes:

[0134] determining the coincidence of the label of the first video in each dimension and the label of the second video in each dimension to obtain the coincidence index corresponding to each dimension;

[0135] The clustering tendency degree corresponding to the video pair is obtained according to the coincidence index corresponding to each dimension.

[0136] In one embodiment, the determination of the coincidence of the label of the first video in each dimension and the label of the second video in each dimension to obtain the coincidence index corresponding to each dimension includes:

[0137] obtaining the coincidence degree of the label of the first video in each dimension and the label of the second video in each dimension; and in response to the case that the coincidence degree is greater than the preset threshold value corresponding to each dimension, determining the coincidence index corresponding to each dimension as a first index;

[0138] The clustering tendency degree corresponding to the video pair is obtained according to the coincidence index corresponding to each dimension, including: the clustering tendency degree corresponding to the video pair is obtained according to the number of the first index.

[0139] In one embodiment, the clustering result is obtained according to the clustering tendency degree corresponding to each video pair, further including:

[0140] Cluster the above at least two videos in descending order of clustering tendency to obtain the above clustering results.

[0141] Furthermore, Figure 12 A schematic diagram of a hardware structure for implementing the method provided in the embodiments of this application is shown. This device can participate in or include the apparatus or system provided in the embodiments of this application. Figure 12 As shown, device 10 may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 12 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, device 10 may also include a... Figure 12 The more or fewer components shown, or having the same Figure 12 The different configurations shown.

[0142] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the device 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0143] The memory 104 can be used to store software programs of application software and modules, such as the program instructions / data storage means corresponding to the method described in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e., implements the above-mentioned video data processing method. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include memories remotely arranged with respect to the processor 102, which can be connected to the device 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0144] The transmission device 106 is used to receive or send data via a network. The specific examples of the above-mentioned network can include a wireless network provided by a communication provider of the device 10. In one example, the transmission device 106 includes a network interface controller (NIC) which can be connected to other network devices through a base station so as to be able to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module which is used to communicate with the Internet in a wireless manner.

[0145] The display can be, for example, a touch screen type liquid crystal display (LCD) which can enable a user to interact with the user interface of the device 10 (or mobile device).

[0146] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above-mentioned embodiments of the present application are described for specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from the order in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.

[0147] Each of the embodiments in the embodiments of the present application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and server embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

[0148] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or can be instructed by programs to complete the related hardware. The above-mentioned programs can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0149] The above merely describes the preferred embodiments of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method of processing video data, the method comprising: The method comprises: acquiring at least two videos; for each of the at least two videos, determining labels of at least two dimensions in the video; for a video pair formed by the at least two videos, acquiring a coincidence degree of a label of each dimension of a first video and a label of each dimension of a second video; in response to a case that the coincidence degree is greater than a preset threshold corresponding to each dimension, determining a coincidence index corresponding to each dimension as a first index; obtaining a clustering tendency degree corresponding to the video pair according to a quantity of the first index, the video pair comprising the first video and the second video, and the clustering tendency degree being used to indicate a probability that the first video and the second video belong to similar videos; obtaining a clustering result according to the clustering tendency degree corresponding to each of the video pairs.

2. The method of claim 1, wherein, The determining, for each of the at least two videos, of the labels of the at least two dimensions in the video comprises: for each of the at least two videos, determining at least two of a label corresponding to a picture dimension, a label corresponding to an audio dimension, and a label corresponding to a text dimension in the video.

3. The method of claim 1 or 2, wherein: in a case that the labels of the at least two dimensions comprise the label corresponding to the picture dimension, the determining of the labels of the at least two dimensions in the video comprises: extracting at least two video frames in the video; and performing picture recognition on each video frame to obtain the label corresponding to the picture dimension; in a case that the labels of the at least two dimensions comprise the label corresponding to the audio dimension, the determining of the labels of the at least two dimensions in the video comprises: performing classification recognition on audio information in the video to obtain the label corresponding to the audio dimension; in a case that the labels of the at least two dimensions comprise the label corresponding to the text dimension, the determining of the labels of the at least two dimensions in the video comprises: extracting text information in the video to obtain the label corresponding to the text dimension.

4. The method of claim 3, wherein, in the case that the labels of the at least two dimensions comprise the label corresponding to the picture dimension, the picture recognition on each video frame to obtain the label corresponding to the picture dimension comprises: performing picture recognition on each video frame to obtain at least one recognition result corresponding to the video frame, the recognition result comprising a picture label and a confidence degree corresponding to the picture label; determining, as a picture label corresponding to the video frame, a picture label whose confidence degree meets a first requirement; counting the picture labels corresponding to the video frames, and determining a label corresponding to the picture dimension according to a picture label whose repetition number meets a second requirement.

5. The method of claim 1, wherein, The obtaining of the clustering result according to the clustering tendency degree corresponding to each of the video pairs comprises: clustering each of the video pairs in a descending order of the clustering tendency degree to obtain the clustering result.

6. A video data processing apparatus, comprising: The apparatus comprises: a video acquisition module configured to acquire at least two videos; a dimension label acquisition module configured to determine, for each of the at least two videos, labels of at least two dimensions in the video; The clustering module is configured to: obtain a coincidence degree of a label of the first video in each dimension and a label of the second video in the each dimension for a video pair formed by the at least two videos; determine a coincidence index corresponding to the each dimension as a first index in response to a case that the coincidence degree is greater than a preset threshold corresponding to the each dimension; obtain a clustering tendency degree corresponding to the video pair according to a quantity of the first indexes; the video pair includes the first video and the second video, and the clustering tendency degree is used to indicate a probability that the first video and the second video belong to similar videos; and obtain a clustering result according to the clustering tendency degrees corresponding to each of the video pairs.

7. A computer readable storage medium characterized in that, The computer readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the video data processing method in any one of claims 1 to 5.

8. An electronic device, comprising: The computer readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the video data processing method in any one of claims 1 to 5.

9. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instruction is executed by the processor to implement the video data processing method in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Label extraction method and device

    CN111222500A

  • Video clustering method and device thereof

    CN113515668A