Video processing and video content searching method and device, electronic equipment and medium
By segmenting the video media asset library into scenes and clustering faces, the problems of low search efficiency and poor accuracy in existing technologies have been solved, achieving efficient and accurate video content search and improving the accuracy of person tracking and clustering results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2022-12-09
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies suffer from low efficiency and poor accuracy when searching for video content related to a specific person from massive video media asset libraries, especially when the person tracking results are inaccurate when the scene changes.
By segmenting videos in the media asset library into scenes, performing character tracking and face clustering, including two rounds of face clustering within and between videos, the video content matching the target face is determined.
It improves search efficiency and accuracy, ensures the accuracy of personnel tracking results, enhances the accuracy of clustering results, and enables fast and accurate historical media data retrieval.
Smart Images

Figure CN116168316B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to video processing and video content retrieval methods, apparatuses, electronic devices and media in the fields of computer vision, big data processing, image processing and databases. Background Technology
[0002] In practical applications, the following situation often occurs: a person in the media asset library suddenly becomes a negative or popular figure. Consequently, it is necessary to search through a massive amount of video assets already in the database to find video content related to or matching that person, in order to modify the historical analysis tag information corresponding to that video content. However, current search methods are usually quite complex to implement, resulting in low search efficiency. Summary of the Invention
[0003] This disclosure provides methods, apparatus, electronic devices, and media for video processing and video content retrieval.
[0004] A video processing method, comprising:
[0005] The video in the media asset library is segmented into scenes, and the people are tracked in the segmented scene segments to obtain the people tracking results.
[0006] Based on the tracking results of people corresponding to the same video, perform face clustering within the same video;
[0007] Based on the face clustering results within the video, face clustering is performed between different videos. This is used to determine the video content in the media asset library that matches the target face, where the target face is the face in the face image to be searched.
[0008] A method for finding video content includes:
[0009] Obtain the face image to be searched, and identify the face in the face image as the target face;
[0010] Based on the face clustering results between different videos, the video content in the media asset library that matches the target face is determined; wherein, the face clustering results between different videos are obtained by performing face clustering between different videos based on the face clustering results within the video, and the face clustering results within the video are obtained by performing face clustering within the same video based on the person tracking results corresponding to the same video, and the person tracking results are obtained by performing scene segmentation on the videos in the media asset library and performing person tracking on the segmented scene segments.
[0011] A video processing apparatus includes: a segmentation and tracking module, a first clustering module, and a second clustering module;
[0012] The segmentation and tracking module is used to segment videos in the media asset library into scenes, track people in the segmented scene segments, and obtain the people tracking results.
[0013] The first clustering module is used to perform face clustering within the same video based on the person tracking results corresponding to the same video;
[0014] The second clustering module is used to perform face clustering between different videos based on the face clustering results within the video, and to determine the video content in the media asset library that matches the target face based on the face clustering results between different videos, wherein the target face is the face in the face image to be searched.
[0015] A video content search device includes: an image acquisition module and a content search module;
[0016] The image acquisition module is used to acquire the face image to be searched and to identify the face in the face image as the target face;
[0017] The content search module is used to determine the video content in the media asset library that matches the target face based on the face clustering results between different videos; wherein, the face clustering results between different videos are obtained by performing face clustering between different videos based on the face clustering results within the video, the face clustering results within the video are obtained by performing face clustering within the same video based on the person tracking results corresponding to the same video, and the person tracking results are obtained by performing scene segmentation on the videos in the media asset library and performing person tracking on the segmented scene segments.
[0018] An electronic device, comprising:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described above.
[0022] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described above.
[0023] A computer program product includes a computer program / instructions that, when executed by a processor, implement the method described above.
[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0025] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0026] Figure 1 This is a flowchart of an embodiment of the video processing method described in this disclosure;
[0027] Figure 2 This is a schematic diagram illustrating the process of scene segmentation and face clustering within a video as described in this disclosure.
[0028] Figure 3 This is a flowchart illustrating an embodiment of the video content search method described in this disclosure;
[0029] Figure 4 This is a schematic diagram of the composition structure of embodiment 400 of the video processing apparatus described in this disclosure;
[0030] Figure 5 This is a schematic diagram of the composition structure of Embodiment 500 of the video content search device described in this disclosure;
[0031] Figure 6 A schematic block diagram of an electronic device 600 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0032] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0033] Furthermore, it should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0034] Figure 1 This is a flowchart illustrating an embodiment of the video processing method described in this disclosure. Figure 1 As shown, the specific implementation methods are as follows.
[0035] In step 101, the video in the media asset library is segmented into scenes, and the segmented scene segments are used for character tracking to obtain character tracking results.
[0036] In step 102, face clustering is performed within the same video based on the person tracking results corresponding to the same video.
[0037] In step 103, based on the face clustering results within the video, face clustering is performed between different videos. This is used to determine the video content in the media asset library that matches the target face based on the face clustering results between different videos. The target face is the face in the face image to be searched.
[0038] By employing the scheme described in the above embodiments, through operations such as scene segmentation, person tracking, and face clustering, the video content matching the face in the face image to be searched can be determined. This reduces complexity and improves search efficiency. Furthermore, scene segmentation improves the accuracy of person tracking results, and two rounds of face clustering—one within the video and one between videos—improve the accuracy of clustering results, thereby enhancing the accuracy of search results.
[0039] In traditional methods, people are usually tracked directly based on media asset stream data. However, the videos in the media asset library often have various scene changes, which can easily lead to inaccurate people tracking results.
[0040] Therefore, the solution described in this disclosure proposes that each video in the media asset library can be segmented into scenes, and that character tracking can be performed on each segmented scene fragment to obtain character tracking results.
[0041] For any video in the media asset library, it can first be split into frames, and then scene segmentation technology can be used to divide it into scene segments. The number of scene segments may be one or more.
[0042] Character tracking can be performed on each segmented scene, such as using a tracking model. Accordingly, during character tracking, forced segmentation can be applied based on the scene segment. For example, in the traditional way, the tracking model tracks character A within the time interval [t1,t2]. However, if t1 to t2 spans a scene, such as [t1,ts] belonging to one scene segment and [ts,t2] belonging to another scene segment, with ts located after t1 and before t2, then the tracking model can be controlled to perform character tracking separately for the two time intervals [t1,ts] and [ts,t2].
[0043] The above processing method incorporates the consideration of shot segmentation, that is, scene segmentation technology is used to segment the scene, thereby improving the accuracy of character tracking results without affecting processing speed.
[0044] Preferably, the person tracking result may include: a first face feature (featureID) and the first tracking result information corresponding to the first face feature, wherein the first face feature is the face feature of the tracked face.
[0045] For a scene segment, there may be one person or multiple people. If there are multiple people, then for each person, the first facial feature and the corresponding first tracking result information can be obtained separately. Here, to distinguish it from other facial features and tracking result information that appear later, the facial feature of the tracked face is called the first facial feature, and the corresponding tracking result information is called the first tracking result information. Similar cases will not be elaborated on later.
[0046] The first tracking result information may include: the video in which the face is located (i.e., the video in which the face is located), and the frame information in the video. For example, the frame information can be marked by the start frame and the end frame. If necessary, other information may also be included, such as the position information of the face in the frame.
[0047] Furthermore, based on the tracking results of people corresponding to the same video, face clustering can be performed within the same video.
[0048] Preferably, for the tracking results of each person corresponding to the same video, the following processing can be performed: cluster the first facial features in each person tracking result according to the similarity to obtain the similarity clustering result, that is, obtain each cluster. For each obtained similarity clustering result, determine the second facial features and second tracking result information corresponding to each similarity clustering result according to the first facial features and the corresponding first tracking result information included therein.
[0049] For example, video a is divided into 10 scene segments, namely scene segment 1 to scene segment 10. Assuming that for each scene segment, a person tracking result is obtained, namely person tracking result 1 to person tracking result 10, then the first facial features in these 10 person tracking results can be clustered according to similarity. Assuming that a total of 3 similarity clustering results are obtained, for these 3 similarity clustering results, the second facial features and second tracking result information corresponding to these 3 similarity clustering results can be determined respectively based on the first facial features and the corresponding first tracking result information included in them.
[0050] The above processing can efficiently and accurately achieve face clustering within the video, thus laying a good foundation for subsequent processing.
[0051] Preferably, for any similarity clustering result, the method for determining its corresponding second face feature and second tracking result information may include: selecting the optimal face from the faces corresponding to the first face feature included in the similarity clustering result, taking the first face feature corresponding to the optimal face as the second face feature corresponding to the similarity clustering result, summarizing the first tracking result information corresponding to the first face feature included in the similarity clustering result, and taking the summarized result as the second tracking result information corresponding to the similarity clustering result.
[0052] For example, if a certain similarity clustering result b corresponding to video a includes three first face features, then the optimal face can be selected from the faces corresponding to these three first face features. The first face feature corresponding to the selected optimal face can be used as the second face feature corresponding to the similarity clustering result b. The first tracking result information corresponding to these three first face features can be summarized, and the summarized result can be used as the second tracking result information corresponding to the similarity clustering result b.
[0053] In other words, identical faces in the same video can be clustered and merged into a single face. The first facial feature of the best face can be retained as an index. For example, the first facial feature corresponding to the best face can be used as the second facial feature corresponding to the similarity clustering result b. In addition, the second tracking result information can include the first tracking result information corresponding to the previous three first facial features. For example, if the face corresponding to the similarity clustering result b appears in frames 10-20, 40-55, and 70-82 of video a, respectively, and belongs to different scene segments, then the second tracking result information corresponding to the similarity clustering result b will record that the face appeared in frames 10-20, 40-55, and 70-82 of video a, respectively.
[0054] Preferably, the method for selecting the optimal face from the faces corresponding to the first facial features included in any similarity clustering result may include: obtaining the scores of the faces corresponding to each first facial feature included in the similarity clustering result according to a pre-set scoring criterion, and taking the face with the highest score as the optimal face.
[0055] The specific criteria used for the predetermined scoring are not limited and can be determined according to actual needs. For example, scoring can be based on the face probability of each face, with a higher face probability resulting in a higher score; or scoring can be based on the blurriness of each face, with a lower blurriness resulting in a higher score; or scoring can be based on the profile angle of each face, with a smaller profile angle resulting in a higher score, and so on.
[0056] By performing optimal face selection, the identified second face features can be made more accurate, thereby improving the accuracy of subsequent processing results.
[0057] Figure 2 This diagram illustrates the process of scene segmentation and face clustering within a video, as described in this disclosure. Figure 2 As shown, the video can first be divided into multiple different scene segments. Then, character tracking can be performed on each scene segment to obtain the character tracking results. Assuming that the faces in the second row of the image (in order from top to bottom) represent the faces tracked from different scene segments, then further, face clustering can be performed within the video. For example, the white faces in the image can be clustered into one category, and the black faces in the image can be clustered into another category, thus obtaining different similarity clustering results, corresponding to white faces and black faces respectively.
[0058] After completing face clustering within each video, face clustering between different videos can also be performed based on the face clustering results within each video.
[0059] Preferably, incremental clustering can be performed on the second face features corresponding to the similarity clustering results of each video in the media asset library to obtain each cluster center feature (clusterID). Then, for each cluster center feature, the following processing can be performed: obtain the second tracking result information corresponding to the second face feature corresponding to the cluster center feature, summarize the obtained second tracking result information, and use the summary result as the third tracking result information corresponding to the cluster center feature, which is used to determine the video content in each video that matches the target face based on each cluster center feature and the corresponding third tracking result information.
[0060] By clustering faces across different videos, the same face in different videos can be merged, and each cluster center feature can correspond to a different face.
[0061] As can be seen, the scheme described in this disclosure adopts a two-stage clustering approach: the first stage is face clustering within a video, and the second stage is face clustering between different videos. This approach minimizes the impact of face angles and blurriness on the clustering results, thereby improving the accuracy of the clustering results.
[0062] Based on the above processing results, actual face search can be performed, thus realizing the function of historical media data retrieval.
[0063] Accordingly, Figure 3 This is a flowchart illustrating an embodiment of the video content search method described in this disclosure. Figure 3 As shown, the specific implementation methods are as follows.
[0064] In step 301, the face image to be searched is obtained, and the face in the face image is identified as the target face.
[0065] In step 302, based on the face clustering results between different videos, the video content in the media asset library that matches the target face is determined; wherein, the face clustering results between different videos are obtained by performing face clustering between different videos based on the face clustering results within the video, the face clustering results within the video are obtained by performing face clustering within the same video based on the person tracking results corresponding to the same video, and the person tracking results are obtained by performing scene segmentation on the videos in the media asset library and performing person tracking on the segmented scene segments.
[0066] By employing the scheme described in the above embodiments, through operations such as scene segmentation, person tracking, and face clustering, the video content matching the face in the face image to be searched can be determined. This reduces complexity and improves search efficiency. Furthermore, scene segmentation improves the accuracy of person tracking results, and two rounds of face clustering—one within the video and one between videos—improve the accuracy of clustering results, thereby enhancing the accuracy of search results.
[0067] Preferably, the face clustering results between different videos may include: features of each cluster center and corresponding third tracking result information, where each cluster center feature corresponds to a different face. Accordingly, the method for determining the video content in each video in the media asset library that matches the target face based on the face clustering results between different videos may include: obtaining the third face features of the target face, filtering out the cluster center features that match the third face features from the cluster center features, using the third tracking result information corresponding to the matching cluster center features as the target tracking result information, and determining the video content in each video in the media asset library that matches the target face based on the target tracking result information.
[0068] There are no restrictions on how to obtain facial features for any given face; for example, existing mature implementation methods can be used.
[0069] Furthermore, preferably, the similarity between each cluster center feature and the third face feature can be obtained separately. The cluster center feature with the highest similarity can then be used as the matching cluster center feature. Alternatively, cluster center features with a similarity greater than a predetermined threshold can be used as the matching cluster center features. The specific value of the threshold can be determined according to actual needs, and the specific method used can also be determined based on actual requirements, making it very flexible and convenient.
[0070] Preferably, the third tracking result information may include: the video where the corresponding face is located, and the frame information of the corresponding face in the video. Accordingly, the face corresponding to the target tracking result information can be identified as the target face, and the video where the target face is located and the frame information of the target face in the video can be determined based on the target tracking result information.
[0071] For example, assuming there is only one matching cluster center feature, the third tracking result information corresponding to that cluster center feature can be obtained as the target tracking result information. If the target tracking result records that the corresponding face appears in video a, video m and video x respectively, and specifically records the frame information in each video, then the corresponding frames in video a, video m and video x can be used as the video content that matches the target face.
[0072] Whether it is the first tracking result information, the second tracking result information, or the third tracking result information, if necessary, it can also include some other information, such as the position information of the face in the frame.
[0073] Furthermore, there are no restrictions on the specific storage format / data structure of the information, whether it is the first tracking result information, the second tracking result information, or the third tracking result information; it can be determined according to actual needs.
[0074] Through the above processing, a rapid historical media asset retrieval function can be achieved, such as finding the required video content from tens of millions of hours of historical media assets at a speed of seconds.
[0075] In addition, for the video content found, the corresponding historical analysis tag information can be modified as needed. That is, the analyzed person information can be modified with one click without re-analyzing the media assets already in the database.
[0076] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this disclosure. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure. Furthermore, for parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0077] In summary, the proposed solution in this disclosure provides a fast database browsing method based on scene segmentation and face clustering. This method can determine the video content to be searched with low implementation complexity, thereby improving search efficiency. Furthermore, scene segmentation improves the accuracy of person tracking results, and two face clustering operations improve the accuracy of clustering results, thereby improving the accuracy of search results.
[0078] The above is an introduction to the method embodiments. The following describes the solution described in this disclosure further through device embodiments.
[0079] Figure 4 This is a schematic diagram of the structural composition of embodiment 400 of the video processing apparatus described in this disclosure. Figure 4 As shown, it may include: a segmentation tracking module 401, a first clustering module 402, and a second clustering module 403.
[0080] The segmentation and tracking module 401 is used to segment videos in the media asset library into scenes, track people in the segmented scene segments, and obtain the people tracking results.
[0081] The first clustering module 402 is used to perform face clustering within the same video based on the person tracking results corresponding to the same video.
[0082] The second clustering module 403 is used to perform face clustering between different videos based on the face clustering results within the video, and to determine the video content in the media asset library that matches the target face based on the face clustering results between different videos, wherein the target face is the face in the face image to be searched.
[0083] By employing the scheme described in the above-mentioned device embodiment, video content matching the face in the face image to be searched can be determined through operations such as scene segmentation, person tracking, and face clustering. This reduces complexity and improves search efficiency. Furthermore, scene segmentation improves the accuracy of person tracking results, and two rounds of face clustering—one within the video and one between videos—improve the accuracy of clustering results, thereby enhancing the accuracy of search results.
[0084] Traditional methods typically involve directly tracking people based on media asset stream data. However, videos in a media asset library often exhibit various scene changes, which can easily lead to inaccurate tracking results. To address this, the solution described in this disclosure proposes that a segmentation tracking module 401 segment each video in the media asset library into scenes, and then perform person tracking on each of the segmented scene segments to obtain the tracking results. For any video in the media asset library, it can first be split into frames, and then scene segmentation technology can be used to divide it into scene segments. The number of scene segments obtained may be one or more.
[0085] Preferably, the person tracking result may include: a first facial feature and first tracking result information corresponding to the first facial feature, wherein the first facial feature is the facial feature of the tracked face.
[0086] Furthermore, based on the tracking results of people corresponding to the same video, face clustering can be performed within the same video.
[0087] Preferably, the first clustering module 402 can perform the following processing on the tracking results of each person corresponding to the same video: cluster the first face features in each person tracking result according to the similarity to obtain the similarity clustering result; and for each obtained similarity clustering result, determine the second face features and second tracking result information corresponding to each similarity clustering result according to the first face features and the corresponding first tracking result information included therein.
[0088] Preferably, for any similarity clustering result, the method by which the first clustering module 402 determines its corresponding second face feature and second tracking result information may include: selecting the optimal face from the faces corresponding to the first face features included in the similarity clustering result, taking the first face feature corresponding to the optimal face as the second face feature corresponding to the similarity clustering result, summarizing the first tracking result information corresponding to the first face features included in the similarity clustering result, and taking the summarized result as the second tracking result information corresponding to the similarity clustering result.
[0089] Preferably, the method by which the first clustering module 402 selects the optimal face from the faces corresponding to the first facial features included in any similarity clustering result may include: obtaining the scores of the faces corresponding to each first facial feature included in the similarity clustering result according to a pre-set scoring criterion, and taking the face with the highest score as the optimal face.
[0090] After completing the face clustering within each video, face clustering between different videos can be performed based on the face clustering results within each video.
[0091] Preferably, the second clustering module 403 can perform incremental clustering on the second face features corresponding to the similarity clustering results of each video in the media asset library to obtain each cluster center feature. Then, for each cluster center feature, the following processing can be performed respectively: obtain the second tracking result information corresponding to the second face feature corresponding to the cluster center feature, summarize the obtained second tracking result information, and use the summary result as the third tracking result information corresponding to the cluster center feature, which is used to determine the video content that matches the target face in each video based on each cluster center feature and the corresponding third tracking result information.
[0092] Figure 5 This is a schematic diagram of the structural composition of Embodiment 500 of the video content search device described in this disclosure. Figure 5 As shown, it may include: an image acquisition module 501 and a content search module 502.
[0093] Image acquisition module 501 is used to acquire the face image to be searched and identify the face in the face image as the target face.
[0094] The content search module 502 is used to determine the video content in the media asset library that matches the target face based on the face clustering results between different videos. The face clustering results between different videos are obtained by performing face clustering between different videos based on the face clustering results within the video. The face clustering results within the video are obtained by performing face clustering within the same video based on the person tracking results corresponding to the same video. The person tracking results are obtained by segmenting the videos in the media asset library into scenes and performing person tracking on the segmented scene segments.
[0095] By employing the scheme described in the above-mentioned device embodiment, video content matching the face in the face image to be searched can be determined through operations such as scene segmentation, person tracking, and face clustering. This reduces complexity and improves search efficiency. Furthermore, scene segmentation improves the accuracy of person tracking results, and two rounds of face clustering—one within the video and one between videos—improve the accuracy of clustering results, thereby enhancing the accuracy of search results.
[0096] Preferably, the face clustering results between different videos may include: features of each cluster center and corresponding third tracking result information, with each cluster center feature corresponding to a different face. Accordingly, the content search module 502 may determine the video content in the media asset library that matches the target face based on the face clustering results between different videos by: obtaining the third face features of the target face, filtering out the cluster center features that match the third face features from the cluster center features, using the third tracking result information corresponding to the matching cluster center features as the target tracking result information, and determining the video content in the media asset library that matches the target face based on the target tracking result information.
[0097] Preferably, the content search module 502 can obtain the similarity between each cluster center feature and the third face feature, and then use the cluster center feature with the highest similarity as the matching cluster center feature, or use the cluster center feature with a similarity greater than a predetermined threshold as the matching cluster center feature.
[0098] In addition, preferably, the third tracking result information may include: the video where the corresponding face is located, and the frame information where the corresponding face is located in the video. Accordingly, the content lookup module 502 can determine the face corresponding to the target tracking result information as the target face, and determine the video where the target face is located and the frame information where it is located in the video based on the target tracking result information.
[0099] Figure 4 and Figure 5 The specific workflow of the device embodiment shown can be found in the relevant descriptions in the foregoing method embodiments, and will not be repeated here.
[0100] In summary, the device embodiments of this disclosure propose a fast database browsing method based on scene segmentation and face clustering. This method can determine the video content to be searched with low implementation complexity, thereby improving search efficiency. Moreover, scene segmentation improves the accuracy of person tracking results, and two face clustering operations improve the accuracy of clustering results, thereby improving the accuracy of search results.
[0101] The solutions described in this disclosure can be applied to the field of artificial intelligence, particularly in areas such as computer vision, big data processing, image processing, and databases. Artificial intelligence is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware and software technologies. Artificial intelligence hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0102] The videos described in the embodiments of this disclosure are not targeted at any specific user and do not reflect the personal information of any specific user. The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solutions of this disclosure all comply with relevant laws and regulations and do not violate public order and good morals.
[0103] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0104] Figure 6 A schematic block diagram of an electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0105] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0106] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0107] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as those described in this disclosure. For example, in some embodiments, the methods described in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the methods described in this disclosure can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the methods described in this disclosure by any other suitable means (e.g., by means of firmware).
[0108] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0112] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0113] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0114] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0115] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A video processing method, comprising: The video in the media asset library is segmented into scenes, and the people are tracked in the segmented scene segments to obtain the people tracking results. The person tracking result includes: a first face feature and the first tracking result information corresponding to the first face feature. The first face feature is the face feature of the tracked face. The first tracking result information includes: the video where the corresponding face is located, and the frame information where the corresponding face is located in the video. The frame information is the frame information marked by the start frame and the end frame. Based on similarity, the first facial features in the tracking results of each person corresponding to the same video are clustered to obtain similarity clustering results; based on the first facial features included in each similarity clustering result and the corresponding first tracking result information, the second facial features and second tracking result information corresponding to each similarity clustering result are determined. Incremental clustering is performed on the second face features corresponding to the similarity clustering results of each video in the media asset library to obtain each cluster center feature. For each cluster center feature, the second tracking result information corresponding to the second face feature corresponding to the cluster center feature is summarized. The summarized result is used as the third tracking result information corresponding to the cluster center feature. This is used to determine the video content in each video in the media asset library that matches the target face based on each cluster center feature and the corresponding third tracking result information. The target face is the face in the face image to be searched.
2. The method according to claim 1, wherein, The second facial features and second tracking result information corresponding to each similarity clustering result are determined as follows: The following processing is performed on each of the obtained similarity clustering results: The optimal face is selected from the faces corresponding to the first face feature included in the similarity clustering result, and the first face feature corresponding to the optimal face is used as the second face feature corresponding to the similarity clustering result. The first tracking result information corresponding to the first face feature included in the similarity clustering result is summarized, and the summarized result is used as the second tracking result information corresponding to the similarity clustering result.
3. The method according to claim 2, wherein, The step of selecting the optimal face from the faces corresponding to the first facial features included in the similarity clustering results includes: According to the pre-set scoring criteria, the scores of the faces corresponding to the first facial features included in the similarity clustering results are obtained respectively, and the face with the highest score is taken as the optimal face.
4. A method for searching video content, comprising: Obtain the face image to be searched, and identify the face in the face image as the target face; Based on the face clustering results between different videos, the video content in the media asset library that matches the target face is determined. The face clustering results between different videos are obtained by performing face clustering between different videos based on the face clustering results within the video. The face clustering results within the same video are obtained by performing face clustering within the same video based on the person tracking results corresponding to the same video. The person tracking results are obtained by segmenting the videos in the media asset library into scenes and performing person tracking on the segmented scene segments. The person tracking results include: a first face feature and first tracking result information corresponding to the first face feature. The first face feature is the face feature of the tracked face. The first tracking result information includes: the video where the corresponding face is located, and the frame information where the corresponding face is located in the video. The frame information is the frame information marked by the start frame and end frame within the same video. The face clustering results include: for each obtained similarity clustering result, based on the first face feature and the corresponding first tracking result information included therein, the second face feature and second tracking result information corresponding to each similarity clustering result are determined. Each similarity clustering result is obtained by clustering the first face feature in the tracking results of each person corresponding to the video based on similarity. The face clustering results between different videos include: each cluster center feature and the corresponding third tracking result information. Each cluster center feature corresponds to a different face. Each cluster center feature is obtained by incrementally clustering the second face feature corresponding to the similarity clustering result of each video in the media asset library. The third tracking result information is a summary result obtained by summarizing the second tracking result information corresponding to the second face feature corresponding to each cluster center feature.
5. The method according to claim 4, wherein, The step of determining the video content in the media asset library that matches the target face based on the face clustering results between different videos includes: Obtain the third facial features of the target face; Cluster center features that match the third face features are selected from the cluster center features; The third tracking result information corresponding to the matching cluster center features is used as the target tracking result information, and the video content that matches the target face in each video in the media asset library is determined based on the target tracking result information.
6. The method according to claim 5, wherein, The step of selecting cluster center features that match the third face feature from each cluster center feature includes: The similarity between the features of each cluster center and the third face features is obtained respectively; The cluster center feature with the highest similarity is used as the matching cluster center feature, or the cluster center feature with a similarity greater than a predetermined threshold is used as the matching cluster center feature.
7. The method according to claim 5 or 6, wherein, The third tracking result information includes: the video where the corresponding face is located, and the frame information of the corresponding face in the video. The step of determining the video content in the media asset library that matches the target face based on the target tracking result information includes: identifying the face corresponding to the target tracking result information as the target face, and determining the video where the target face is located and the frame information of the video where it is located based on the target tracking result information.
8. A video processing apparatus, comprising: The module includes a segmentation and tracking module, a first clustering module, and a second clustering module. The segmentation and tracking module is used to segment videos in the media asset library into scenes, track people in the segmented scene segments, and obtain people tracking results. The people tracking results include: a first face feature and first tracking result information corresponding to the first face feature. The first face feature is the face feature of the tracked face. The first tracking result information includes: the video where the corresponding face is located, and the frame information where the corresponding face is located in the video. The frame information is the frame information marked by the start frame and the end frame. The first clustering module is used to cluster the first facial features in the tracking results of each person corresponding to the same video based on similarity, and obtain similarity clustering results; and to determine the second facial features and second tracking results information corresponding to each similarity clustering result based on the first facial features included in each similarity clustering result and the corresponding first tracking result information. The second clustering module is used to incrementally cluster the second face features corresponding to the similarity clustering results of each video in the media asset library to obtain each cluster center feature. For each cluster center feature, the second tracking result information corresponding to the second face feature corresponding to the cluster center feature is summarized, and the summarized result is used as the third tracking result information corresponding to the cluster center feature. This module is used to determine the video content in each video in the media asset library that matches the target face based on each cluster center feature and the corresponding third tracking result information. The target face is the face in the face image to be searched.
9. The apparatus according to claim 8, wherein, The first clustering module performs the following processing on each obtained similarity clustering result: selects the optimal face from the faces corresponding to the first face features included in the similarity clustering results, uses the first face features corresponding to the optimal face as the second face features corresponding to the similarity clustering results, summarizes the first tracking result information corresponding to the first face features included in the similarity clustering results, and uses the summary result as the second tracking result information corresponding to the similarity clustering results.
10. The apparatus according to claim 9, wherein, The first clustering module obtains the scores of the faces corresponding to the first facial features included in the similarity clustering results according to the pre-set scoring criteria, and takes the face with the highest score as the optimal face.
11. A video content retrieval device, comprising: Image acquisition module and content search module; The image acquisition module is used to acquire the face image to be searched and to identify the face in the face image as the target face; The content search module is used to determine the video content in the media asset library that matches the target face based on the face clustering results between different videos. The face clustering results between different videos are obtained by performing face clustering between different videos based on the face clustering results within the video. The face clustering results within the same video are obtained by performing face clustering within the same video based on the person tracking results corresponding to the same video. The person tracking results are obtained by segmenting the videos in the media asset library into scenes and performing person tracking on the segmented scene segments. The person tracking results include: a first face feature and first tracking result information corresponding to the first face feature. The first face feature is the face feature of the tracked face. The first tracking result information includes: the video where the corresponding face is located, and the frame information where the corresponding face is located in the video. The frame information is the frame information marked by the start frame and end frame. The face clustering results within the same video include: for each obtained similarity clustering result, based on the first face feature and the corresponding first tracking result information included therein, the second face feature and second tracking result information corresponding to each similarity clustering result are determined. The similarity clustering results are obtained by clustering the first face feature in the tracking results of each person corresponding to the video based on similarity. The face clustering results between different videos include: each cluster center feature and the corresponding third tracking result information. Each cluster center feature corresponds to a different face. The cluster center feature is obtained by incrementally clustering the second face feature corresponding to the similarity clustering results of each video in the media asset library. The third tracking result information is a summary result obtained by summarizing the second tracking result information corresponding to the second face feature corresponding to each cluster center feature.
12. The apparatus according to claim 11, wherein, The content search module obtains the third facial features of the target face, filters out the cluster center features that match the third facial features from each cluster center feature, uses the third tracking result information corresponding to the matching cluster center feature as the target tracking result information, and determines the video content in each video in the media asset library that matches the target face based on the target tracking result information.
13. The apparatus according to claim 12, wherein, The content search module obtains the similarity between each cluster center feature and the third face feature, and uses the cluster center feature with the highest similarity as the matching cluster center feature, or uses the cluster center feature with a similarity greater than a predetermined threshold as the matching cluster center feature.
14. The apparatus according to claim 12 or 13, wherein, The third tracking result information includes: the video where the corresponding face is located, and the frame information of the corresponding face in the video. The content search module identifies the face corresponding to the target tracking result information as the target face, and determines the video where the target face is located and the frame information within the video based on the target tracking result information.
15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.
17. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the method of any one of claims 1-7.
Citation Information
Patent Citations
Face clustering based video categorization method and retrieval method as well as systems thereof
CN103530652A
Method and device for establishing face index, processing server and storage medium
CN110543584A