Scene Clustering Method, Device and Related Equipment in Video

Through the deep learning model, a multi-frame image in the video is classified and featured, and a data set of scenic spot clustering features is generated and clustered analysis is performed, which solves the problem of low accuracy in video scene recognition and realizes accurate recognition of scenic spot images of different angles or degrees of exposure.

CN114299435BActive Publication Date: 2025-07-18BEIJING IQIYI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111649894.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-07-18
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

In the prior art, video scene recognition has the problem of low recognition accuracy, especially when identifying scenic spot images, image information from all angles of the scenic spot cannot be accurately obtained, resulting in frequent identification errors.

Method used

By obtaining multi-frame images in the video, using deep learning models for classification recognition and feature extraction, obtaining scenic spot images and performing scene classification marks, generating scenic spot clustering feature data sets, and finally performing clustering analysis to obtain clustering results of each scene classification label.

Benefits of technology

It improves the accuracy of the recognition of attractions in the video, and can accurately identify attractions with different angles or degrees of exposure as the same type of attractions, improving the accuracy and consistency of the recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299435B_ABST
    Figure CN114299435B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method for scene clustering in a video, including: obtaining multiple frames of images in the video; performing classification and recognition on the multiple frames of images to obtain scenic spot images in the multiple frames of images; performing scene classification and marking on the scenic spot images according to scene classification labels to obtain the marked scenic spot images; extracting features from the marked scenic spot images to obtain a scenic spot clustering feature dataset; performing clustering analysis based on the scenic spot clustering feature dataset to obtain clustering results corresponding to each scene classification label. In the embodiment of the present invention, after obtaining multiple frames of images in the video, after marking the multiple frames of images, the images are input into a deep learning model for processing to obtain the clustering results corresponding to the marks. According to the clustering results, two scenic spot images with different angles or different exposure degrees in the same type of scenic spot images can be accurately recognized as the same type of scenic spot images, achieving the effect of improving the accuracy of recognizing scene pictures by obtaining the clustering results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image recognition technology, and in particular to a method and apparatus for scene clustering in a video and related devices. Background Art

[0002] During the filming of film and television dramas, well-known scenic spots or internet-famous locations are usually used to improve the filming effect of the whole drama. Therefore, during the viewing process, the audience will also have a need to know the specific location information of some of the filming locations.

[0003] However, there is currently a large error in scene recognition. Selecting individual scenic spot images with relatively high representativeness for recognition will result in too few recognized images that can be obtained during the video playback process, affecting the recognition process. In addition, when the recognition model obtains the scenic spot images, it generally cannot accurately obtain the image information of all angles of the scenic spot, resulting in the inability to accurately recognize the scenic spot or the occurrence of recognition errors during the recognition process, and there is a problem of low recognition accuracy in scene recognition. Summary of the Invention

[0004] A method, apparatus and related devices for scene clustering in a video provided by an embodiment of the present invention solve the problem of low recognition accuracy in scene recognition in the prior art.

[0005] In a first aspect, an embodiment of the present invention provides a method for scene clustering in a video, including:

[0006] Obtain multiple frames of images in the video;

[0007] Perform classification recognition on the multiple frames of images to obtain the scenic spot images in the multiple frames of images;

[0008] Perform scene classification marking on the scenic spot images according to scene classification labels to obtain the marked scenic spot images;

[0009] Extract features from the marked scenic spot images to obtain a scenic spot clustering feature data set;

[0010] Perform clustering analysis based on the scenic spot clustering feature data set to obtain clustering results corresponding to each scene classification label.

[0011] Optionally, the performing classification recognition on the multiple frames of images to obtain the scenic spot images in the multiple frames of images includes:

[0012] Input the multiple frames of images into a pre-trained first deep learning model for classification recognition to obtain the scenic spot images in the multiple frames of images.

[0013] Optionally, before inputting the multi-frame images into the pre-trained first deep learning model for classification and recognition to obtain the scenic spot images in the multi-frame images, the following steps are further included:

[0014] Obtain the created classification model;

[0015] Train the classification model with preset training samples, where the training samples include first scenic spot sample images and first non-scenic spot sample images;

[0016] Determine the trained classification model as the first deep learning model.

[0017] Optionally, the feature extraction of the marked scenic spot images to obtain the scenic spot clustering feature dataset includes:

[0018] Input the marked scenic spot images into the pre-trained second deep learning model for feature extraction to obtain the scenic spot clustering feature dataset.

[0019] Optionally, before inputting the marked scenic spot images into the pre-trained second deep learning model for feature extraction to obtain the scenic spot clustering feature dataset, the following steps are further included:

[0020] Obtain the created feature extraction model;

[0021] Train the feature extraction model with sample images, where the sample images are generated by performing image augmentation on the second scenic spot sample images;

[0022] Determine the trained feature extraction model as the second deep learning model.

[0023] Optionally, the training of the feature extraction model with sample images, where the sample images are generated by performing image processing on the second scenic spot sample images, includes:

[0024] Input the sample images into the feature extraction model to extract sample features;

[0025] Generate a scene classification feature library based on the sample features;

[0026] Train the feature extraction model to obtain residual network parameters according to the scene classification feature library and the classification function, where the classification function is generated based on the landmark feature library;

[0027] Update the feature extraction model based on the residual network.

[0028] Optionally, the clustering analysis based on the scenic spot clustering feature dataset to obtain the clustering results corresponding to each scene classification label includes:

[0029] Obtain multiple scene classification clustering clusters based on the scenic spot clustering feature dataset, where the scene classification clustering clusters match the scene classification labels;

[0030] Calculate the correlation between any two of the multiple scene classification clustering clusters to obtain a correlation value, where the any two scene classification clustering clusters have the same scene classification label;

[0031] If the correlation value is less than or equal to a preset threshold, merge the two scene classification clustering clusters into a new scene classification clustering cluster, where the new scene classification clustering cluster includes at least two of the scene classification labels;

[0032] Repeat the correlation calculation for any two of the scene classification clustering clusters until the correlation values of any two of the scene classification clustering clusters are all greater than the preset threshold, and obtain the clustering results corresponding to each scene classification label.

[0033] In a second aspect, an embodiment of the present invention further provides a scene clustering device in a video, including:

[0034] An acquisition module, configured to acquire multiple frames of images in a video;

[0035] An identification module, configured to perform classification and identification on the multiple frames of images to obtain scenic spot images in the multiple frames of images;

[0036] A classification module, configured to perform scene classification marking on the scenic spot images according to scene classification labels to obtain the marked scenic spot images;

[0037] An extraction module, configured to extract features from the marked scenic spot images to obtain a scenic spot clustering feature dataset;

[0038] An analysis module, configured to perform clustering analysis based on the scenic spot clustering feature dataset to obtain the clustering results corresponding to each scene classification label.

[0039] In a third aspect, an embodiment of the present invention further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor, where when the program or instruction is executed by the processor, the steps of the scene clustering method in a video as described in any one of the above are implemented.

[0040] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, characterized in that a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the scene clustering method in a video as described in any one of the above are implemented.

[0041] An embodiment of the present invention provides a method, apparatus, and related device for scene clustering in a video. The method includes: obtaining multiple frames of images in the video; classifying and identifying the multiple frames of images to obtain scenic images in the multiple frames of images; performing scene classification and marking on the scenic images according to scene classification labels to obtain marked scenic images; extracting features from the marked scenic images to obtain a scenic clustering feature dataset; and performing clustering analysis based on the scenic clustering feature dataset to obtain clustering results corresponding to each scene classification label. The method for scene clustering in a video provided by the embodiment of the present invention, after obtaining multiple frames of images in the video, performs marking on the multiple frames of images, and then inputs the images into a deep learning model for processing to obtain clustering results corresponding to the marks. According to the clustering results, two scenic images with different angles or different exposure degrees in the same type of scenic images can be accurately identified as the same type of scenic images, achieving the effect of improving the accuracy of identifying scene pictures by obtaining the clustering results. Description of the Drawings

[0042] Figure 1 It is a flowchart of a method for scene clustering in a video according to an embodiment of the present invention;

[0043] Figure 2 It is a schematic structural diagram of scenic classification according to an embodiment of the present invention;

[0044] Figure 3 It is a schematic flowchart of a method for scene clustering in a video according to an embodiment of the present invention;

[0045] Figure 4 It is a schematic structural diagram of a device for scene clustering in a video according to an embodiment of the present invention;

[0046] Figure 5 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed Embodiments

[0047] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0048] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0049] In addition, terms such as "first", "second", etc. may be used herein to describe various directions, actions, steps, or elements, etc., but these directions, actions, steps, or elements are not limited by these terms. These terms are only used to distinguish the first direction, action, step, or element from another direction, action, step, or element. For example, without departing from the scope of the present application, the first speed difference can be the second speed difference, and similarly, the second speed difference can be referred to as the first speed difference. Both the first speed difference and the second speed difference are speed differences, but they are not the same speed difference. The terms "first", "second", etc. should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0050] Figure 1 The method flowchart of a method for scene clustering in a video provided for the embodiments of the present invention. A method for scene clustering in a video provided in this embodiment includes:

[0051] Step 110: Obtain multiple frames of images in the video.

[0052] In this embodiment, the video is a video during the user's viewing process. Specifically, the video includes various scenic images and non-scenic images. Among them, the scenic images include human landscapes or natural sceneries, etc., such as a captured image of Tiananmen. For the distinction between scenic images and non-scenic images, it is mainly based on whether the feature with a relatively large proportion in the image belongs to a scenic spot or a non-scenic spot. Exemplarily, when an image includes both a scenic spot and a pedestrian, if the pedestrian accounts for a relatively small proportion at this time, the image is recognized as a scenic image.

[0053] After processing the video, multiple frames of images in the video are obtained. These multiple frames of images include both scenic images and non-scenic images.

[0054] Step 120: Classify and identify the multiple frames of images to obtain the scenic images in the multiple frames of images.

[0055] Refer to Figure 2 , Figure 2 which is a schematic structural diagram of scenic spot classification in an embodiment of the present invention. Specifically, the input images can be divided into scenic spot images and non-scenic spot images. In this embodiment, multiple frames of images are classified and recognized to distinguish the scenic spot images from the non-scenic spot images in the multiple frames of images. Specifically, the multiple frames of images in multiple videos can be classified and recognized through a deep learning model or other recognition models. According to the results of the classification and recognition, it can be obtained which of the multiple frames of images belong to scenic spot images and which belong to non-scenic spot images. Among them, non-scenic spot images are usually close-ups, such as images of a person or an object as the main body, and scenic spot images are images of buildings or landscapes. And the scenic spot images can be further divided into different types of scenic spots such as skyscrapers, pavilions, and commercial streets.

[0056] Step 130: Perform scene classification and marking on the scenic spot images according to the scene classification labels to obtain the marked scenic spot images.

[0057] In this embodiment, the scene classification labels are relevant labels for distinguishing different scenic spots. Specifically, the labels can be skyscrapers, commercial streets, pavilions, etc. Generally, the relevant labels of the scenic spot images are identified manually or by machine and the scenic spot images are marked. After classifying and marking different scenic spot images, the marked scenic spot images are obtained. Exemplarily, for example, if the captured scenic spot image is the Oriental Pearl Tower, the label marked for it is skyscraper.

[0058] Step 140: Extract features from the marked scenic spot images to obtain a scenic spot clustering feature dataset.

[0059] In this embodiment, features are extracted from the marked scenic spot images and a scenic spot clustering feature dataset is generated according to the features. Specifically, the scenic spot clustering feature dataset includes scenic spot images of the same kind with a relatively high degree of similarity, such as the same building or landscape. Specifically, the marked scenic spot images can be input into a deep learning model for feature extraction. The scenic spot clustering feature dataset represents the feature data of the same scenic spot. Specifically, the similarity between features can be used as a measure of the similarity of scenic spots. Exemplarily, when any captured image of a scenic spot is obtained, if the extracted features are the same as the features in the scenic spot clustering feature dataset, it can be considered that the captured image of the scenic spot belongs to the same scenic spot as the scenic spot corresponding to the scenic spot clustering feature dataset.

[0060] Step 150: Perform clustering analysis based on the scenic spot clustering feature dataset to obtain clustering results corresponding to each scene classification label.

[0061] In this embodiment, by using a scene clustering method based on hierarchical clustering to perform clustering analysis on the scenic spot clustering feature dataset, the clustering results of each scene classification label are obtained, and the tight clusters are merged layer by layer according to certain conditions in a bottom-up manner. Specifically, the clustering results include scenic spots of the same type with different angles and exposure degrees. In the subsequent recognition process, when encountering images of the same type of scenic spots with different angles, the two images can be accurately recognized as the same type of scenic spots through the clustering results. After obtaining the clustering results, when the user needs to recognize a new scenic spot image, by inputting the new scenic spot image into the recognition model containing the clustering results, the recognition model can identify whether the new scenic spot image belongs to the scenic spot images already included in the clustering results. If so, the new scenic spot image is classified as the scenic spot image already included in the clustering results. Exemplarily, in practical applications, for example, by using the recognition model to identify the side view and front view of the Oriental Pearl in the clustering of the TV tower, it can be calculated that the similarity between the front view and the side view is high and they belong to the same cluster. Therefore, both the side view and the front view are recognized as the Oriental Pearl and used as the same clustering result. When other views of the TV tower similar to the Oriental Pearl are input into the recognition model later, the recognition model can also recognize them as TV towers according to the clustering results, thus achieving the effect of improving the accuracy of recognizing scene pictures.

[0062] Specifically, the clustering results can merge similar scenes, thereby achieving temporal connection. For example, in the case of a large time span, the same building may have characteristic changes, and through clustering analysis, it can be recognized as the same building, improving the temporal consistency of the video recognition results and enhancing the user experience.

[0063] A scene clustering method in a video provided by an embodiment of the present invention, after obtaining multiple frames of images in the video, marking the multiple frames of images, and then inputting the images into a deep learning model for processing to obtain the clustering results corresponding to the marks. According to the clustering results, two scenic spot images with different angles or exposure degrees in the same type of scenic spot images can be accurately recognized as the same type of scenic spot images, achieving the effect of improving the accuracy of recognizing scene pictures by obtaining the clustering results.

[0064] In another embodiment, optionally, the classifying and recognizing the multiple frames of images to obtain the scenic spot images in the multiple frames of images includes:

[0065] Inputting the multiple frames of images into a pre-trained first deep learning model for classification recognition to obtain the scenic spot images in the multiple frames of images.

[0066] In this embodiment, the pre-trained first deep learning model is a trained convolutional network image. A pre-trained convolutional network image classification model is used to classify and identify the pushed images. Specifically, according to the classification results, it is possible to obtain which of the multiple frames of images belong to scenic images and which belong to non-scenic images. Among them, non-scenic images are usually close-ups, for example, images whose main body may be a person or an object. In this embodiment, the scenic classification method can be any common image classification method, including but not limited to methods based on deep learning algorithms.

[0067] Optionally, before inputting the multiple frames of images into the pre-trained first deep learning model for classification and identification to obtain the scenic images in the multiple frames of images, it further includes:

[0068] Obtain the created classification model;

[0069] Train the classification model with a preset training sample, where the training sample includes a first scenic sample image and a first non-scenic sample image;

[0070] Determine the trained classification model as the first deep learning model.

[0071] In this embodiment, first obtain the established classification model. This classification model can be a deep learning model that conforms to image classification, and no specific limitation is made in this embodiment. A suitable model can be selected according to the actual situation. The preset training sample is a large number of sample images, which include a large number of first scenic sample images and first non-scenic sample images. Input the large number of first scenic sample images and first non-scenic sample images into the classification model for training to obtain and continuously adjust the various parameters of the classification model. Finally, after the classification model is trained, it can directly identify whether any frame of the video is a scenic image or a non-scenic image.

[0072] Optionally, the feature extraction of the marked scenic images to obtain the scenic clustering feature dataset includes:

[0073] Input the marked scenic images into the pre-trained second deep learning model for feature extraction to obtain the scenic clustering feature dataset.

[0074] In this embodiment, the pre-trained second deep learning model is a trained convolutional neural network, and the marked scenic spot images are input into the trained convolutional neural network to extract features of the scenic spot images, thereby obtaining a scenic spot clustering feature data set. The scenic spot clustering feature data set represents the feature data of the same scenic spot. Specifically, the similarity between the features can be used as a measure of the similarity of the scenic spots. The scenic spot clustering feature data set is used to divide the same scenic spot images into the same category of images. For example, when a photographed image of any scenic spot is obtained, if the extracted features are the same as the features in the scenic spot clustering feature data set, it can be considered that the photographed image of any scenic spot belongs to the same scenic spot as the scenic spot corresponding to the scenic spot clustering feature data set.

[0075] Optionally, before inputting the marked scenic spot image into a pre-trained second deep learning model for feature extraction and obtaining the scenic spot cluster feature data set, the method further includes:

[0076] Get the created feature extraction model;

[0077] Training the feature extraction model through a sample image, wherein the sample image is generated after image processing based on a sample image of a second scenic spot;

[0078] The trained feature extraction model is determined as the second deep learning model.

[0079] In this embodiment, a feature extraction model is first established. The feature extraction model can be a deep residual network or other deep learning network. Specifically, a deep residual network trained in a public data set is used to extract features from the target image and then clustering is performed. A large number of sample images can be obtained from a public landmark database. Specifically, data augmentation can be performed on the basis of the public landmark database. For example, the input image is randomly cropped or cut out to deliberately create information missing. In this way, the model is prompted to fill in the missing parts from the global information, so that the model avoids excessive focus on local information, improves the global information extraction ability and generalization of the model, and enables the model to correctly identify the same sample under occlusion, human interference and multi-angle transformation. The feature extraction model is trained by a large number of sample images, and the parameters in the model are continuously updated. Finally, a trained feature extraction model is obtained and the feature extraction model is determined as the second deep learning model.

[0080] Optionally, the training of the feature extraction model by using a sample image, wherein the sample image is generated after image augmentation based on the second scenic spot sample image, comprises:

[0081] Inputting the sample image into the feature extraction model to extract sample features;

[0082] Generate a scene classification feature library based on the sample features;

[0083] Train the feature extraction model according to the scene classification feature library and the classification function to obtain the residual network parameters, where the classification function is generated based on the landmark feature library;

[0084] Update the feature extraction model based on the residual network.

[0085] In this embodiment, a pre-trained deep residual network on a public dataset is used to extract features from the target image; and density-based clustering algorithm (Density-Based Spatial Clustering of Applications with Noise, abbreviated as DBSCAN) clustering is performed on the features corresponding to all training data; according to the clustering labels, each category represents a landmark sub-cluster, and representative features are selected from all the features within the landmark sub-cluster and stored in a dictionary, thereby establishing a landmark feature library. The network is trained based on the dictionary labels of the landmark feature library through a classification function, thereby updating the residual network parameters, where the classification function can use the softmax classification function. Finally, with the fixed trained parameters, the network inputs the target region image and obtains the landmark clustering features. Exemplarily, the training dataset is the Google Open Landmark Recognition System Google-Landmarks-v2, which has 200,000 landmarks and 4 million image instances. At the beginning of the training stage, the model parameters pre-trained on the ImageNet visualization database are used for initialization, and features are extracted from the training set images, and then clustering is performed based on the feature data. Here, the density-based clustering method is used for clustering, and other similar unsupervised clustering methods can also be used. Then, using the clustered landmark ID as the key and the average value of all features under the clustering center as the representative feature value, the network is backpropagated by setting the contrast loss function, and the residual model parameters are updated in the way of momentum update. The core of this self-paced contrast learning training framework is the clustering-based pseudo-label algorithm, which uses the clustering labels as supervision information and realizes network update in the form of a contrast loss function.

[0086] Among them, finally, when the model converges, the scenic spot recognition model has the ability to distinguish different landmark scenes. When the image detection data of different scenic spots is input, the similarity between features can measure the similarity of scenic spots. After hierarchical clustering from bottom to top based on this feature, the same scenic spot features can obtain the same labels, thereby obtaining the scenic spot clustering result.

[0087] Optionally, the clustering analysis based on the scenic spot clustering feature dataset to obtain the clustering results corresponding to the scene classification labels includes:

[0088] Obtain multiple scene classification clustering clusters based on the scenic spot clustering feature dataset, and the scene classification clustering clusters match the scene classification labels;

[0089] Calculate the correlation between any two scene classification clustering clusters among the multiple scene classification clustering clusters to obtain a correlation value;

[0090] If the correlation value is less than or equal to a preset threshold, merge the two scene classification clustering clusters into a new scene classification clustering cluster, and the new scene classification clustering cluster includes at least two of the scene classification labels;

[0091] Repeat the correlation calculation for any two scene classification clustering clusters until the correlation values of any two scene classification clustering clusters are both greater than the preset threshold to obtain the clustering results corresponding to each scene classification label.

[0092] In this embodiment, hierarchical clustering analysis is performed through the scenic spot clustering feature dataset, that is, clustering processing is performed again. Specifically, hierarchical clustering is a bottom-up method of hierarchically merging tight clusters according to certain conditions. Based on the scenic spot features, hierarchical clustering is performed on the scenic spot images. For the situation where the scenic spot scenes in the video are the same but have different angles and exposure degrees, they are clustered into a unified category to provide information support for subsequent recognition. Specifically, the scenic spot clustering feature dataset is divided, and each individual scenic spot image corresponds to a separate scenic spot clustering cluster, and the relevant features of the scenic spot are included in the scenic spot clustering cluster. Specifically, the clustering results can merge similar scenes, thereby achieving temporal connection. For example, in the case of a large time span, the same building may have changing features, and through clustering analysis, it can be recognized as the same building, improving the temporal consistency of the video recognition results and enhancing the user experience.

[0093] Specifically, each scenic spot image is regarded as a new clustering cluster; the correlation calculation is to calculate the average distance of the squared distances between each pair of elements included between every two clustering clusters, and merge the two clustering clusters with a distance less than the threshold; if the distance is greater than this threshold, the two clustering clusters are separated, where the threshold is set to 0.5. Specifically, this threshold can be adaptively adjusted according to the actual situation, and in this embodiment, 0.5 is taken as an example for illustration. Repeat the correlation calculation for any two scenic spot clustering clusters until all clustering clusters are merged to obtain the first-level clustering result. Specifically, in this embodiment, only the clustering clusters with the same label can be merged. Exemplarily, for example, even if the similarity between two clustering clusters with the labels of skyscrapers and pavilions is relatively high, they cannot be merged. In this embodiment, the second-level clustering is performed by combining the scene classification label and the first-level clustering result. Specifically, the features corresponding to each scenic spot image cluster in the first-level clustering result are averaged and aggregated as the representative of this clustering cluster, and the similarity of the scenic spot clustering features between clusters is calculated. Combining the scenic spot label information such as skyscrapers, pavilions, commercial streets, etc., when the similarity between two clusters is greater than the threshold and the label information is the same, they are merged into a new clustering cluster, otherwise they are not merged. This threshold can be adaptively adjusted according to the actual situation and is not specifically limited in this embodiment. After traversing all the clustering clusters, the final scene clustering result is obtained. A scene clustering method in a video provided by an embodiment of the present invention, after obtaining multiple frames of images in the video, marking the multiple frames of images, and then inputting the images into a deep learning model for processing to obtain the clustering result corresponding to the mark. According to the clustering result, two scenic spot images with different angles or different exposure degrees in the same type of scenic spot images can be accurately recognized as the same type of scenic spot images, achieving the effect of improving the accuracy of recognizing scene pictures by obtaining the clustering result.

[0094] Refer to Figure 3 , Figure 3The following is a schematic flowchart of a method for scene clustering in a video in this embodiment. First, relevant video frames in the video are obtained, and the video frames are classified and recognized for scenic spots through a backbone network for scenic spot classification. If the video frame is recognized as a secondary classification after classification (i.e., scenic spot images of different types, such as skyscrapers, pavilions, etc.), subsequent secondary clustering processing is performed. If the video frame cannot be directly recognized, it is determined whether the video frame is a scenic spot image through primary classification. If it is a non-scenic spot image, the video frame is discarded. If it is a scenic spot image, the scenic spot image is subjected to feature extraction to identify features, and scenic spot images with similar features are subjected to primary clustering (HAC clustering) to obtain a primary clustering result. The primary clustering result is subjected to secondary clustering through labels (i.e., scenic spot images of different types, such as skyscrapers, pavilions, etc.) to obtain a clustering result. This clustering result can identify the same type of scenic spots including different angles and different degrees of exposure. In subsequent recognition processes, when encountering images of the same type of scenic spots with different angles, the two images can be accurately recognized as the same type of scenic spots through the clustering result.

[0095] Figure 4 The following is a schematic structural diagram of a scene clustering device 200 in a video provided in this embodiment. The scene clustering device 200 provided in this embodiment includes:

[0096] An acquisition module 210, configured to acquire multiple frames of images in the video;

[0097] An identification module 220, configured to perform classification and recognition on the multiple frames of images to obtain scenic spot images in the multiple frames of images;

[0098] A classification module 230, configured to perform scene classification marking on the scenic spot images according to scene classification labels to obtain marked scenic spot images;

[0099] An extraction module 240, configured to perform feature extraction on the marked scenic spot images to obtain a scenic spot clustering feature data set;

[0100] An analysis module 250, configured to perform clustering analysis based on the scenic spot clustering feature data set to obtain clustering results corresponding to each scene classification label.

[0101] Optionally, the identification module 220 includes:

[0102] An identification sub-module, configured to input the multiple frames of images into a pre-trained first deep learning model for classification and recognition to obtain scenic spot images in the multiple frames of images.

[0103] Optionally, it further includes:

[0104] A first creation module, configured to acquire a created classification model;

[0105] A first training module for training the classification model with preset training samples, where the training samples include first scenic spot sample images and first non-scenic spot sample images;

[0106] A first determination module for determining the trained classification model as the first deep learning model.

[0107] Optionally, the extraction module 240 includes:

[0108] An extraction sub-module for inputting the marked scenic spot images into a pre-trained second deep learning model for feature extraction to obtain a scenic spot clustering feature dataset.

[0109] Optionally, it further includes:

[0110] A second creation module for obtaining the created feature extraction model;

[0111] A second training module for training the feature extraction model with sample images, where the sample images are generated after image augmentation based on second scenic spot sample images;

[0112] A second determination module for determining the trained feature extraction model as the second deep learning model.

[0113] Optionally, the second training module includes:

[0114] A feature extraction sub-module for inputting the sample images into the feature extraction model to extract sample features;

[0115] A feature generation sub-module for generating a scene classification feature library based on the sample features;

[0116] A model training sub-module for training the feature extraction model according to the scene classification feature library and a classification function to obtain residual network parameters, where the classification function is generated based on the landmark feature library;

[0117] Updating the feature extraction model based on the residual network.

[0118] Optionally, the analysis module 250 includes:

[0119] An acquisition sub-module for acquiring multiple scene classification clustering clusters based on the scenic spot clustering feature dataset, where the scene classification clustering clusters match the scene classification labels;

[0120] A calculation sub-module for calculating the correlation between any two of the multiple scene classification clustering clusters to obtain a correlation value, where any two of the scene classification clustering clusters have the same scene classification label;

[0121] A merging sub-module, configured to merge the two scene classification clusters into a new scene classification cluster if the correlation value is less than or equal to a preset threshold, where the new scene classification cluster includes at least two of the scene classification labels;

[0122] A generating sub-module, configured to repeatedly calculate the correlation between any two scene classification clusters until the correlation values between any two scene classification clusters are all greater than the preset threshold, so as to obtain the clustering results corresponding to each scene classification label.

[0123] A scene clustering device in a video provided by an embodiment of the present invention, after obtaining multiple frames of images in the video, marking the multiple frames of images, and then inputting the images into a deep learning model for processing to obtain the clustering results corresponding to the marks, realizes that according to the clustering results, two scenic spot images with different angles or different exposure degrees in the same type of scenic spot images can be accurately recognized as the same type of scenic spot images, and realizes the effect of improving the accuracy of recognizing scene pictures by obtaining the clustering results.

[0124] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention, as Figure 5 shown, the electronic device includes a memory 310 and a processor 320. The number of processors 320 in the electronic device 300 may be one or more. Figure 5 Taking one processor 320 as an example; the memory 310 and the processor 320 in the server may be connected through a bus or other means. Figure 5 Taking the connection through the bus as an example.

[0125] The memory 310, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the scene clustering method in a video in an embodiment of the present invention. The processor 320 runs the software programs, instructions, and modules stored in the memory 310, thereby performing various functional applications and data processing of the server / terminal / server, that is, implementing the above-mentioned scene clustering method in a video.

[0126] Among them, the processor 320 is used to run the computer program stored in the memory 310 to implement the following steps:

[0127] Obtain multiple frames of images in the video;

[0128] Perform classification recognition on the multiple frames of images to obtain the scenic spot images in the multiple frames of images;

[0129] Perform scene classification marking on the scenic spot images according to scene classification labels to obtain the marked scenic spot images;

[0130] Extract features from the marked scenic spot images to obtain a scenic spot clustering feature dataset;

[0131] Perform clustering analysis based on the scenic spot clustering feature dataset to obtain clustering results corresponding to each scene classification label.

[0132] Optionally, the classifying and recognizing the multi-frame images to obtain the scenic spot images in the multi-frame images includes:

[0133] Input the multi-frame images into a pre-trained first deep learning model for classification and recognition to obtain the scenic spot images in the multi-frame images.

[0134] Optionally, before inputting the multi-frame images into a pre-trained first deep learning model for classification and recognition to obtain the scenic spot images in the multi-frame images, it further includes:

[0135] Obtain the created classification model;

[0136] Train the classification model with preset training samples, where the training samples include first scenic spot sample images and first non-scenic spot sample images;

[0137] Determine the trained classification model as the first deep learning model.

[0138] Optionally, the extracting features from the marked scenic spot images to obtain a scenic spot clustering feature dataset includes:

[0139] Input the marked scenic spot images into a pre-trained second deep learning model for feature extraction to obtain a scenic spot clustering feature dataset.

[0140] Optionally, before inputting the marked scenic spot images into a pre-trained second deep learning model for feature extraction to obtain a scenic spot clustering feature dataset, it further includes:

[0141] Obtain the created feature extraction model;

[0142] Train the feature extraction model with sample images, where the sample images are generated after image processing based on second scenic spot sample images;

[0143] Determine the trained feature extraction model as the second deep learning model.

[0144] Optionally, the training the feature extraction model with sample images, where the sample images are generated after image augmentation based on the second scenic spot sample images includes:

[0145] Input the sample images into the feature extraction model to extract sample features;

[0146] Generate a scene classification feature library based on the sample features, where the scene classification feature library includes the scene classification labels;

[0147] Train the feature extraction model according to the scene classification feature library and a classification function to obtain residual network parameters, where the classification function is generated based on the landmark feature library;

[0148] Update the feature extraction model based on the residual network.

[0149] Optionally, the clustering analysis based on the scenic spot clustering feature dataset to obtain the clustering results corresponding to each scene classification label includes:

[0150] Obtain multiple scene classification clustering clusters based on the scenic spot clustering feature dataset, where the scene classification clustering clusters match the scene classification labels;

[0151] Calculate the correlation between any two of the multiple scene classification clustering clusters to obtain a correlation value, where any two of the scene classification clustering clusters have the same scene classification label;

[0152] If the correlation value is less than or equal to a preset threshold, merge the two scene classification clustering clusters into a new scene classification clustering cluster, where the new scene classification clustering cluster includes at least two of the scene classification labels;

[0153] Repeat the correlation calculation for any two of the scene classification clustering clusters until the correlation values of any two of the scene classification clustering clusters are all greater than the preset threshold to obtain the clustering results corresponding to each scene classification label.

[0154] In one embodiment, for an electronic device provided by an embodiment of the present invention, the computer program is not limited to the above method operations, and can also execute related operations in the scene clustering method in the video provided by any embodiment of the present invention.

[0155] The memory 310 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 310 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 310 may further include a memory remotely set relative to the processor 320, and these remote memories may be connected to the server / terminal / server through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0156] An electronic device for scene clustering in a video provided by an embodiment of the present invention, after obtaining multiple frames of images in the video, marking the multiple frames of images, and then inputting the images into a deep learning model for processing to obtain a clustering result corresponding to the mark, realizes that according to the clustering result, two scenic spot images with different angles or different exposure degrees in the same type of scenic spot images can be accurately recognized as the same type of scenic spot images, and realizes the effect of improving the accuracy of recognizing scene pictures by obtaining the clustering result.

[0157] An embodiment of the present invention also provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a method for scene clustering in a video when executed by a computer processor. The method includes:

[0158] Obtain multiple frames of images in the video;

[0159] Perform classification and recognition on the multiple frames of images to obtain scenic spot images in the multiple frames of images;

[0160] Perform scene classification marking on the scenic spot images according to scene classification labels to obtain marked scenic spot images;

[0161] Extract features from the marked scenic spot images to obtain a scenic spot clustering feature dataset;

[0162] Perform clustering analysis based on the scenic spot clustering feature dataset to obtain clustering results corresponding to each scene classification label.

[0163] Of course, the computer-executable instructions of a storage medium containing computer-executable instructions provided by an embodiment of the present invention are not limited to the method operations as described above, and can also execute related operations in a method for scene clustering in a video provided by any embodiment of the present invention.

[0164] The computer-readable storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0165] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0166] The program code contained on the storage medium may be transmitted by any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.

[0167] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or terminal. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through the Internet service provider via the Internet).

[0168] A storage medium for scene clustering in a video provided by an embodiment of the present invention, after obtaining multiple frames of images in the video, marking the multiple frames of images, and then inputting the images into a deep learning model for processing to obtain a clustering result corresponding to the marking, realizes that according to the clustering result, two scenic spot images with different angles or different exposure degrees in the same type of scenic spot images can be accurately recognized as the same type of scenic spot images, and realizes the effect of improving the accuracy of recognizing scene pictures by obtaining the clustering result.

[0169] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for scene clustering in a video, characterized in that, Including: Obtain multiple frames of images in the video; Classify and identify the multiple frames of images to obtain scenic spot images in the multiple frames of images; Perform scene classification marking on the scenic spot images according to scene classification labels to obtain the marked scenic spot images; Extract features from the marked scenic spot images to obtain a scenic spot clustering feature dataset; Perform clustering analysis based on the scenic spot clustering feature dataset to obtain clustering results corresponding to each scene classification label; The performing clustering analysis based on the scenic spot clustering feature dataset to obtain clustering results corresponding to each scene classification label includes: obtaining multiple scene classification clustering clusters based on the scenic spot clustering feature dataset, where the scene classification clustering clusters match the scene classification labels; calculating the correlation between any two of the multiple scene classification clustering clusters to obtain a correlation value, where the any two scene classification clustering clusters have the same scene classification label; if the correlation value is less than or equal to a preset threshold, then merge the two scene classification clustering clusters into a new scene classification clustering cluster, and the new scene classification clustering cluster includes at least two of the scene classification labels; repeat the calculation of the correlation between any two scene classification clustering clusters until the correlation values of any two scene classification clustering clusters are all greater than the preset threshold to obtain the clustering results corresponding to each scene classification label.

2. The method according to claim 1, wherein The classifying and identifying the multiple frames of images to obtain the scenic spot images in the multiple frames of images includes: Input the multiple frames of images into a pre-trained first deep learning model for classification and identification to obtain the scenic spot images in the multiple frames of images.

3. The method according to claim 2, wherein Before inputting the multiple frames of images into a pre-trained first deep learning model for classification and identification to obtain the scenic spot images in the multiple frames of images, it further includes: Obtain a created classification model; Train the classification model with preset training samples, where the training samples include first scenic spot sample images and first non-scenic spot sample images; Determine the trained classification model as the first deep learning model.

4. The method according to claim 1, characterized in that, The extracting features from the marked scenic spot images to obtain a scenic spot clustering feature dataset includes: Input the marked scenic spot images into a pre-trained second deep learning model for feature extraction to obtain a scenic spot clustering feature dataset.

5. The method according to claim 4, characterized in that Before inputting the marked scenic spot images into a pre-trained second deep learning model for feature extraction to obtain a scenic spot clustering feature dataset, it further includes: Obtain a created feature extraction model; Train the feature extraction model with sample images, where the sample images are generated after image augmentation based on second scenic spot sample images; Determine the trained feature extraction model as the second deep learning model.

6. The method according to claim 5, characterized in that, The training the feature extraction model with sample images, where the sample images are generated after image processing based on the second scenic spot sample images includes: Input the sample images into the feature extraction model to extract sample features; Generate a scene classification feature library based on the sample features; Training the feature extraction model according to the scene classification feature library and the classification function to obtain residual network parameters, where the classification function is generated based on a landmark feature library, and the establishment process of the landmark feature library includes: using a pre-trained deep residual network on a public dataset to extract features from a target image; clustering the features corresponding to all training data using a density-based clustering algorithm; according to the clustering labels, each category represents a landmark sub-cluster, and selecting representative features from all features within the landmark sub-cluster and storing them in a dictionary to establish the landmark feature library; Updating the feature extraction model based on the residual network.

7. A scene clustering device in a video, characterized in that, Including: An acquisition module for acquiring multiple frames of images in a video; An identification module for classifying and identifying the multiple frames of images to obtain scenic images in the multiple frames of images; A classification module for performing scene classification marking on the scenic images according to scene classification labels to obtain marked scenic images; An extraction module for extracting features from the marked scenic images to obtain a scenic clustering feature dataset; An analysis module for performing clustering analysis based on the scenic clustering feature dataset to obtain clustering results corresponding to each scene classification label; The analysis module includes: an acquisition sub-module for obtaining multiple scene classification clustering clusters based on the scenic clustering feature dataset, where the scene classification clustering clusters match the scene classification labels; a calculation sub-module for calculating the correlation between any two of the multiple scene classification clustering clusters to obtain a correlation value, where any two of the scene classification clustering clusters have the same scene classification label; a merging sub-module for merging the two scene classification clustering clusters into a new scene classification clustering cluster if the correlation value is less than or equal to a preset threshold, and the new scene classification clustering cluster includes at least two of the scene classification labels; a generation sub-module for repeating the calculation of the correlation between any two scene classification clustering clusters until the correlation values of any two scene classification clustering clusters are all greater than the preset threshold to obtain clustering results corresponding to each scene classification label.

8. An electronic device, characterized in that, Including a processor, a memory, and a program or instruction stored on the memory and executable on the processor, where the program or instruction, when executed by the processor, implements the steps of the method for scene clustering in a video according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, it implements the steps of the method for scene clustering in a video according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Scene classification method and device based on artificial intelligence and electronic equipment

    CN112949620A