Method and device for generating summary of learning video, electronic device and storage medium

CN115713709BActive Publication Date: 2026-09-08BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211339560.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-26
Publication Date
2026-09-08
Estimated Expiration
2042-10-26

AI Technical Summary

Technical Problem

一般来说,用户需要大致观看之后才能了解到视频中是否有感兴趣或符合自身学习基础的内容,因此,选择学习视频需要耗费一定的时间和精力

Benefits of technology

[0019] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods provided in any embodiment of this disclosure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115713709B_ABST
    Figure CN115713709B_ABST
Patent Text Reader

Abstract

The present disclosure provides a learning video summary generation method and device, electronic equipment and storage medium, relates to the technical field of artificial intelligence, in particular to image processing, and can be applied to the online education scene. The specific implementation scheme is: obtaining a first image set based on a learning video; extracting the learning content features of each image in the first image set, and clustering based on the learning content features of each image to obtain a plurality of target image clusters; selecting a corresponding representative image in each target image cluster in the plurality of target image clusters; and obtaining summary information of the learning video based on the representative image corresponding to each target image cluster. The embodiment of the present disclosure is beneficial to quickly understand the learning video and reduce the time and effort cost of selecting the learning video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and in particular to image processing technology, which can be applied to online education scenarios. Background Technology

[0002] With the rise of online education, learning through online videos has become a popular choice for many. Generally, users need to watch a video to determine if it contains content that interests them or aligns with their learning level; therefore, selecting learning videos requires a certain amount of time and effort. Summary of the Invention

[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for generating summaries of learning videos.

[0004] According to one aspect of this disclosure, a method for generating summaries of learning videos is provided, comprising:

[0005] Based on the learning video, a first set of images is obtained;

[0006] The learned content features of each image in the first image set are extracted, and clustering is performed based on the learned content features of each image to obtain multiple target image clusters;

[0007] Select a representative image from each of the multiple target image clusters;

[0008] Based on the representative image corresponding to each target image cluster, the summary information of the learning video is obtained.

[0009] According to another aspect of this disclosure, a summary generation apparatus for learning videos is provided, comprising:

[0010] The frame extraction module is used to obtain the first set of images based on the learning video;

[0011] The image clustering module is used to extract the learned content features of each image in the first image set, and to cluster based on the learned content features of each image to obtain multiple target image clusters;

[0012] The image selection module is used to select a representative image from each of the multiple target image clusters.

[0013] The summary determination module is used to obtain summary information of the learning video based on the representative image corresponding to each target image cluster.

[0014] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods provided in any embodiment of this disclosure.

[0018] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods provided in any embodiment of this disclosure.

[0019] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods provided in any embodiment of this disclosure.

[0020] According to the technical solution of this disclosure, since multiple target image clusters are obtained by clustering the learning content features of each image in the first image set corresponding to the learning video, a representative image of a target image cluster can better represent the learning content in that cluster, and representative images of different target image clusters can contain different learning content. Thus, obtaining summary information of the learning video based on the representative images of each target image cluster allows the summary information to represent the rich learning content in the learning video with a smaller amount of data, thereby facilitating a quick understanding of the learning video and reducing the time and effort cost of selecting a learning video.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0022] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0023] Figure 1 This is a flowchart illustrating a method for generating summaries of learning videos according to an embodiment of this disclosure.

[0024] Figure 2 This is a flowchart illustrating a method for generating summaries of learning videos according to another embodiment of this disclosure.

[0025] Figure 3 This is a schematic diagram of a scenario illustrating a method for generating summaries of learning videos according to an embodiment of this disclosure.

[0026] Figure 4 This is a schematic diagram illustrating an application example of the learning video summary generation method in this embodiment of the present disclosure.

[0027] Figure 5 This is a schematic block diagram of a learning video summary generation apparatus according to an embodiment of the present disclosure.

[0028] Figure 6 This is a schematic block diagram of a learning video summary generation apparatus according to another embodiment of the present disclosure.

[0029] Figure 7 This is a schematic block diagram of a learning video summary generation apparatus according to yet another embodiment of the present disclosure.

[0030] Figure 8 This is a schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0031] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0032] Figure 1 A flowchart illustrating a method for generating summaries of learning videos according to an embodiment of this disclosure is shown. This method can be applied to electronic devices. Figure 1 As shown, the method may include the following steps S110 to S140.

[0033] Step S110: Based on the learning video, obtain the first image set.

[0034] In this embodiment of the disclosure, the learning video may also be referred to as the teaching video. Exemplarily, the learning video may include videos published on an online learning platform or video publishing platform for user self-study. For example, the learning video may be a pre-recorded video of a teacher explaining courseware, including courseware information and audio information of the teacher's explanation.

[0035] In this embodiment of the disclosure, the first image set may include multiple video frames from the learning video, which may be all or some of the video frames in the learning video. That is, the first image set can be composed of each video frame in the learning video, or the video frames in the learning video can be extracted or filtered to obtain the first image set.

[0036] For example, frames can be extracted from the learning video based on a preset frame extraction interval or frame extraction frequency to obtain a first set of images. For instance, a preset frame extraction command can be used to extract frames from the learning video, where the preset frame extraction command may include a frame extraction frequency, a video source address, and a video frame storage address. Executing this command will extract frames from the learning video at the video source address according to the frame extraction frequency, and save the resulting first set of images to the video frame storage address. The frame extraction frequency could be, for example, one frame every 5 seconds, one frame every 3 seconds, etc.

[0037] Step S120: Extract the learning content features of each image in the first image set, and perform clustering based on the learning content features of each image to obtain multiple target image clusters.

[0038] For example, in this embodiment of the disclosure, the learning content features may include features of the document content used for teaching, such as features of text and / or illustrations in courseware shown in a learning video.

[0039] Alternatively, there are multiple ways to extract features from the learning content, as shown in the following examples.

[0040] Example 1: The learning content features of each image in the first image set are extracted using feature point detection algorithms and feature point descriptors. This involves detecting relevant feature points in the image and describing these feature points using a pre-defined feature point descriptor to extract the learning content features. Feature point detection algorithms include, for example, SIFT (Scale-invariant feature transform), SURF (Speeded Up Robust Features), and ORB (Object Request Broker).

[0041] Example 2: A CNN (Convolutional Neural Network) is used to extract the learned content features of each image in the first set. The network architecture used in the CNN includes, for example, AlexNet and ResNet (Residual Network).

[0042] Optionally, when the learning content features include multiple features, the same feature extraction method can be used for all types of features, or different feature extraction methods can be used for different types of features. This disclosure does not limit this.

[0043] For example, the clustering in this embodiment is used to group multiple images with similarity into the same cluster, and different images with large differences in similarity into different clusters. In practical applications, the similarity between images can be determined using the learned content features of the images, thereby determining whether the images belong to the same cluster based on the similarity between them.

[0044] Specifically, for a set of images to be clustered (e.g., a first set of images), a cluster can be created for the first image in the set as the current cluster, and then the other images in the set can be traversed sequentially. Based on the learned features of the traversed images and the learned features of the images in the current cluster, the similarity between the traversed images and the images in the current cluster is determined. If the similarity is greater than or equal to a preset threshold, the traversed image is added to the current cluster. If the similarity is less than the preset threshold, a new cluster is created for the traversed image, and this new cluster is used as the current cluster. It can be understood that multiple clusters can be obtained after the traversal. An image can be selected from the current cluster to calculate the similarity with the traversed images. The selection method can be the first image, the last image in the cluster, etc., depending on the actual needs.

[0045] For example, for the image set {image 1, image 2, image 3, image 4}, first, a cluster 1 containing image 1 is created as the current cluster. Then, the similarity between image 2 and image 1 is determined. If the similarity is greater than or equal to a preset threshold, image 2 is added to cluster 1, so cluster 1 = {image 1, image 2}. Next, the similarity between image 3 and image 1 is determined. If the similarity is less than a preset threshold, a cluster 2 containing image 3 is created, and cluster 2 is used as the current cluster. Finally, the similarity between image 3 and image 4 is determined. If the similarity is greater than or equal to a preset threshold, image 4 is added to cluster 2, so cluster 2 = {image 3, image 4}. Thus, two clusters are obtained.

[0046] For example, the first image set can be clustered based on the learned content features, and the resulting clusters can be used as multiple target image clusters. Alternatively, multiple clustering operations can be performed based on the learned content features and the first image set to obtain target image clusters. For instance, a first clustering operation based on the learned content features can be performed to obtain multiple first image clusters. Then, one image can be selected from each first image cluster to obtain a second image set. A second clustering operation based on the learned content features can be performed to obtain multiple second image clusters, and so on. Multiple clustering operations can be performed, and the final multiple image clusters can be used as multiple target image clusters. Each clustering operation can use the same type of learned content features or different types of learned content features; this disclosure does not limit this.

[0047] Step S130: Select the corresponding representative image from each of the multiple target image clusters.

[0048] For example, a representative image is used to represent each image in its respective image cluster. Optionally, the first or last image in the image cluster can be selected as the representative image, or a preset keyframe extraction algorithm can be used to select keyframes in the image cluster as representative images.

[0049] One exemplary implementation is that the representative image can be the image used during the clustering process to calculate similarity with each image in the image cluster. For example, in the clustering process described above, the first image of the current cluster is used to calculate the similarity with the traversed images, so the representative image can be the first image in the image cluster. In this way, the similarity between the representative image and each image in the image cluster can be less than a preset threshold.

[0050] Step S140: Based on the representative image corresponding to each target image cluster, obtain the summary information of the learning video.

[0051] For example, the summary information can be a collection of images, an image sequence, or courseware or PDF (Portable Document Format) files formed based on the image sequence.

[0052] Optionally, multiple representative images corresponding one-to-one with multiple target image clusters can be combined to obtain summary information. Alternatively, multiple preferred images can be selected from the multiple representative images and combined to obtain summary information. Or, some images that do not meet the conditions can be filtered from the multiple representative images, and the remaining images can be combined to form summary information.

[0053] According to the above method, since multiple target image clusters are obtained by clustering the learning content features of each image in the first image set corresponding to the learning video, the representative image of a target image cluster can effectively represent the learning content in that cluster, and the representative images of different target image clusters can contain different learning content. Thus, obtaining the summary information of the learning video based on the representative images of each target image cluster allows this summary information to represent the rich learning content of the learning video with a smaller amount of data, thereby facilitating a quick understanding of the learning video and reducing the time and effort required for selecting a learning video.

[0054] As described above, in this embodiment of the disclosure, the learning content features may include a variety of features and may be extracted in a variety of ways. For example, the learning content features may include local text features extracted based on feature point detection algorithms, and / or global texture features extracted based on CNNs.

[0055] As an optional implementation, step S120 above, extracting the learning content features of each image in the first image set, may include: extracting local text features of each image in the first image set based on a feature point detection algorithm.

[0056] For example, local text features may include multiple text feature points, where each text feature point can be represented using a preset feature point descriptor operator. For instance, a preset feature point descriptor operator can be used to determine the vector corresponding to the text feature point, which can represent the text feature point.

[0057] For example, these multiple text feature points can include detailed features of the text characterized by low-level image information such as scale, orientation, and size. Since the main content of learning videos often includes text, and text generally has obvious stroke features with relatively rich feature points, feature extraction using feature point detection algorithms and feature point descriptors can achieve the acquisition of multiple salient features in the image using low-level algorithms. This not only enhances the richness and saliency of the learning content features but also improves the efficiency of feature extraction.

[0058] Accordingly, when the learning content features include the aforementioned local text features, the first image set can be clustered based on the local text features to obtain multiple target image clusters.

[0059] As an optional implementation, step S120 above, extracting the learning content features of each image in the first image set, may include: extracting the global texture features of each image in the first image set based on CNN.

[0060] For example, global texture features may include texture variations and color features of illustrations in an image, which can be represented using high-order semantic information. Since CNNs often employ a stacked structure design, where higher-level structures can capture features with larger receptive fields, extracting global texture features through CNNs can improve the accuracy of global texture features and avoid redundant features generated by using low-order features.

[0061] Accordingly, when the learning content features include the aforementioned global texture features, the first image set can be clustered based on the global texture features to obtain multiple target image clusters.

[0062] It should be noted that the above implementation methods can also be combined, that is, in step S120, both local text features and global texture features are extracted. The steps of extracting local text features and extracting global texture features can be performed sequentially or in parallel. If the above combination of local text features and global texture features is used, the learning content features can be extracted from both low-level local features and high-level global features, thereby extracting more comprehensive and accurate features from the image, while also achieving high feature extraction efficiency.

[0063] Accordingly, when the learning content features include local text features and global texture features, clustering can be performed sequentially based on the two features, thereby improving the clustering effect through multi-level clustering and thus enhancing the accuracy of the summary information of the learning video.

[0064] Figure 2 A flowchart illustrating a method for generating summaries of learning videos according to another embodiment of this disclosure is shown. As an optional implementation, such as... Figure 2 As shown, in step S120 above, clustering is performed based on the learned content features of each image to obtain multiple target image clusters, including:

[0065] Step S210: Cluster the first image set based on the local text features of each image to obtain multiple first image clusters;

[0066] Step S220: Based on multiple first image clusters, obtain multiple target image clusters.

[0067] Multiple first image clusters can be used as multiple target image clusters for extracting summary information. Alternatively, multiple first image clusters can be further processed to obtain multiple target image clusters. For example, adjacent first image clusters with a high degree of similarity can be merged to obtain multiple target image clusters; or, multiple first image clusters can be re-clustered to obtain multiple target image clusters.

[0068] The above implementation fully utilizes local textual features in the learning video, thereby clustering images with similar text into the same category and separating images with different text into different categories. In this way, summarizing the learning video based on representative images of each target image cluster allows the summary information to represent the rich textual content of the learning video with a smaller amount of data, thus facilitating a quick understanding of the learning video and reducing the time and effort required for selecting it.

[0069] In an optional implementation, step S121, clustering the first image set based on the local text features of each image to obtain multiple first image clusters, may include:

[0070] Based on the local text features of each image in the first image set and a preset first similarity threshold, the first image set is clustered to obtain multiple second image clusters;

[0071] In each of the multiple second image clusters, a representative image is selected, and a second image set is obtained based on the representative image of each second image cluster;

[0072] Based on the local text features of each image in the second image set and a preset second similarity threshold, the second image set is clustered to obtain multiple first image clusters; wherein, the second similarity threshold is less than the first similarity threshold.

[0073] According to the above implementation method, two clustering operations are performed based on local text features to obtain multiple first image clusters. The first clustering uses a relatively high first similarity threshold, meaning that only images with very high similarity in the first image set will cluster into the same second image cluster. This ensures that when selecting representative images from the second image clusters to form the second image set, only images with very high similarity are filtered out, guaranteeing the number of images in the second image set, thus guaranteeing the recall rate of video frames in the learning video. Correspondingly, the second clustering uses a relatively low second similarity threshold. Therefore, selecting representative images based on the results of the second clustering to obtain the summary information of the learning video can reduce the content redundancy in the summary information. In other words, according to the above implementation method, a balance can be achieved between ensuring recall rate and reducing content redundancy, thereby optimizing the selection effect of summary information.

[0074] As explained above, during clustering, whether two images belong to the same cluster can be determined based on whether the similarity between them exceeds a preset similarity threshold. When clustering using local text features, since local text features include multiple text feature points, RANSAC (Random Sample Consensus) can be used to determine the similarity between two images. Specifically, RANSAC can be used to determine the number of inliners and outliners in the two sets of feature points corresponding to the two images, thereby determining the similarity based on the number of inliners and outliners.

[0075] In an optional implementation, step S122, obtaining multiple target image clusters based on multiple first image clusters, may include:

[0076] A representative image is selected from each of the multiple first image clusters, and a third image set is obtained based on the representative image of each first image cluster;

[0077] Based on the global texture features of each image in the third image set, the third image set is clustered to obtain multiple target image clusters.

[0078] Optionally, the method for selecting a representative image in each first image cluster can be to select the first image, the last image, or a key image in the first image cluster.

[0079] According to the above implementation method, clustering is first performed based on local text features, and then clustering is performed based on global texture features. This multi-level clustering based on multiple features improves the representativeness of the representative images of each image cluster in the clustering results from multiple aspects such as feature accuracy, feature comprehensiveness, image recall rate, and image repetition, and reduces the total amount of data of the representative images, thereby optimizing the extraction effect of summary information.

[0080] In the clustering process based on global texture features described above, it is also possible to determine whether two images belong to the same cluster based on whether the similarity between them is greater than a preset third similarity threshold. The cosine distance between the global texture features of two images can be used as the similarity between them. The global texture features can be represented by vectors.

[0081] It is understandable that in practical applications, clustering can be performed first based on global texture features, and then based on local text features. For example, the first image set can be clustered based on global texture features, then the fourth image set can be determined based on the clustering results, and the fourth image set can be clustered based on local text features to obtain multiple target image clusters. The specific implementation process can be similarly set up with reference to the above implementation method, and will not be elaborated here.

[0082] Figure 3 This illustration shows a scenario diagram of a method for generating summaries of learning videos according to another embodiment of this disclosure. The summary information is courseware in a predetermined format. Figure 3 As shown, the method for generating summaries of the learning video can be executed by electronic device 31. Electronic device 31 can be a server. Electronic device 31 is connected to user device 32, and the method can include the following steps S310 to S350:

[0083] Step S310: Based on the learning video, obtain the first image set.

[0084] Step S320: Extract the learning content features of each image in the first image set, and perform clustering based on the learning content features of each image to obtain multiple target image clusters.

[0085] Step S330: Select the corresponding representative image from each of the multiple target image clusters.

[0086] Step S340: Based on the representative image corresponding to each target image cluster, obtain the summary information of the learning video.

[0087] Steps S310 to S340 are similar to steps S110 to S140 in the aforementioned embodiments and can be implemented with reference to the aforementioned embodiments, and will not be described in detail here.

[0088] Step S350: Output learning videos and courseware to user device 32 so that the courseware and learning videos can be displayed together on the video recommendation page of the user device.

[0089] For example, the user equipment 32 may be a terminal device such as a personal computer, smartphone, or tablet computer. Figure 3 As shown, user device 32 can display multiple different learning videos and associate the summary information of each learning video with that video. In this way, users can browse the courseware associated with each learning video on the video recommendation page to understand the learning content of each video, thereby reducing the cost of selecting learning videos.

[0090] It should be noted that the application of the learning video summary generation method in this embodiment is not limited to this. For example, the method in this embodiment can be used for video deduplication. Specifically, summary information of two videos can be generated, and the similarity between the two videos can be determined based on the similarity between the summary information of the two videos. Thus, when the similarity between the two videos is greater than a threshold, the two videos are determined to be duplicate videos.

[0091] The following provides a specific application example of the method according to embodiments of this disclosure. In this application example, the method for generating summaries of learning videos includes the following steps one through five.

[0092] Step 1: Enter the video source address of the learning video.

[0093] Step 2: Extract frames from the learning video using a preset frame extraction command. This command includes the extraction frequency, video source address, and video frame save address. Executing this command extracts frames from the learning video at the specified extraction frequency and saves the resulting first image set to the specified video frame save address. The structure of the first image set is, for example, {image 1, image 2, image 3, ..., image N}.

[0094] Step 3: Based on the first image set in the video frame storage address, apply a multi-level video aggregation method to obtain multiple target image clusters. The structure of these multiple target image clusters is, for example, {{Image 1, Image 6}, {Image 7, Image 10, ..., Image 14}, ..., {Image N}}. The outermost set contains F inner sets, representing F target image clusters, where F is an integer greater than or equal to 1. Each inner set contains one or more image / video frames.

[0095] The multi-level video aggregation method mainly includes two stages.

[0096] Feature extraction stage:

[0097] Considering that the content in the learning videos mainly consists of text and illustrations, feature extraction is designed from two dimensions: low-level local features and high-level global features, used to extract local text features and global texture features, respectively. Text features are mostly stroke features, and the direction features of strokes are also relatively rich; therefore, feature point detection algorithms and feature point descriptors are used for feature extraction. Illustrations have rich textures and colors and generally contain rich semantic information. Using low-level feature descriptions would generate a lot of redundant features and have poor robustness; therefore, a CNN network structure is used to extract high-level semantic features. Due to its stacked structure, CNNs often have a large receptive field for high-level features, enabling them to capture more global features of the image.

[0098] For each image in the first image set, local text features P can be extracted. l and global texture features F h Among them, local text features P l This includes the features of feature points, and the features of each feature point can be represented by a vector F. l describe.

[0099] Content aggregation stage:

[0100] The input information for the content aggregation stage includes: a first image set {image 1, image 2, image 3, ..., image N}, and the local text features P of each image. l Global texture features F h .

[0101] Multi-level clustering is used in the content aggregation stage. Figure 4 The diagram illustrates this multi-level clustering process, which includes the following three levels of clustering.

[0102] First level: Based on local text features P l Cluster the first set of images.

[0103] This can be achieved using a pre-built clustering algorithm module. This module uses a local text feature similarity calculation module to calculate the similarity between images. It determines whether images belong to the same cluster based on whether the similarity exceeds a similarity threshold. A relatively high similarity threshold is set to ensure high recall of the video content. For example, a similarity threshold of 0.9.

[0104] Suppose that the local text features of image 1 are P l1 The local text features of image 2 are P l2 The processing flow of the above local text feature similarity calculation module is as follows:

[0105] Based on P l1 A set of feature points and P l2 A set of feature points is used to calculate the number N of inner group points using the RANSAC function. I The number of outliers N o ;

[0106] Calculate the matching score, Score, where Score = N I / (N I +N o );

[0107] If Score ≥ Threshold, the two images are determined to belong to the same cluster; otherwise, they are determined not to belong to the same cluster. Threshold is the similarity threshold.

[0108] like Figure 4 As shown, by clustering the first image set, M second image clusters can be output, namely second image cluster 1 to second image cluster M, where M is an integer greater than or equal to 2.

[0109] Based on the above output, the first frame of each second image cluster is used to construct the second image set for clustering at the second level. (See reference) Figure 4 For example, the second set of images is, for instance, {image 1, image 3, image 6, ..., image N}.

[0110] Second level: Based on local text features P l Cluster the second set of images.

[0111] This can be achieved using a pre-built clustering algorithm module, employing a local text feature similarity calculation module to calculate the similarity between images. The specific algorithm principle can be found in the first level. A relatively low similarity threshold is set to reduce the repetition of video content; for example, a similarity threshold of 0.8.

[0112] like Figure 4As shown, by clustering the second image set, K first image clusters can be output, namely first image cluster 1 to first image cluster K, where K is an integer greater than or equal to 2 and less than or equal to M.

[0113] Based on the above output, the first frame of each first image cluster is used to construct a third image set for clustering at the third level. (See reference) Figure 4 For example, the third set of images is, for instance, {image 1, image 6, ..., image N}.

[0114] Third level: Based on global texture features F h Cluster the third set of images.

[0115] This can be achieved using a pre-built clustering algorithm module. The specific algorithm principle can be found in the first level, the difference being the use of a global texture feature similarity calculation module to calculate the similarity between images. The similarity threshold can be set according to actual needs, for example, to 0.99.

[0116] Assume the global texture features of image 1 are F h1 (x1, y1), the global texture feature of image 6 is F h2 (x2, y2), the processing flow of the global texture feature similarity calculation module is as follows:

[0117] Calculate the cosine distance cosθ, where,

[0118] If cosθ ≥ Threshold, the two images are determined to belong to the same cluster; otherwise, they are determined not to belong to the same cluster. Threshold is the similarity threshold.

[0119] like Figure 4 As shown, by clustering the third image set, F target image clusters can be output, namely target image cluster 1 to target image cluster F, where F is an integer greater than or equal to 2 and less than or equal to K.

[0120] Step 4: Select the first frame of each target image cluster as the representative image based on the video time. The selection strategy is not unique; other frames can also be selected as representative images according to actual needs.

[0121] Step 5: Obtain summary information based on each representative image. The structure of this summary information is, for example, {image 1, image 7, ..., image N}.

[0122] As can be seen from the above method, since multiple target image clusters are obtained by clustering the learning content features of each image in the first image set corresponding to the learning video, the representative image of a target image cluster can effectively represent the learning content within that cluster, and the representative images of different target image clusters can contain different learning content. Thus, obtaining the summary information of the learning video based on the representative images of each target image cluster allows this summary information to represent the rich learning content of the learning video with a smaller amount of data, thereby facilitating a quick understanding of the learning video and reducing the time and effort required for selecting it.

[0123] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0124] According to embodiments of this disclosure, this disclosure also provides a summary generation apparatus for learning videos that implement the above-described methods. Figure 5 A schematic block diagram of a learning video summary generation apparatus provided in one embodiment of the present disclosure is shown.

[0125] The frame extraction module 510 is used to obtain a first set of images based on the learning video;

[0126] The image clustering module 520 is used to extract the learning content features of each image in the first image set, and to perform clustering based on the learning content features of each image to obtain multiple target image clusters;

[0127] The image selection module 530 is used to select a corresponding representative image from each of the plurality of target image clusters;

[0128] The abstract determination module 540 is used to obtain the abstract information of the learning video based on the representative image corresponding to each target image cluster.

[0129] In some embodiments, the learning content features include local text features extracted based on a feature point detection algorithm, and / or global texture features extracted based on a convolutional neural network.

[0130] In some embodiments, Figure 5 On the basis of, such as Figure 6 As shown, the image clustering module 530 includes:

[0131] The first clustering unit 610 is used to cluster the first image set based on the local text features of each image to obtain multiple first image clusters;

[0132] The second clustering unit 620 is used to obtain the multiple target image clusters based on the multiple first image clusters.

[0133] In some embodiments, the first clustering unit 610 is used to:

[0134] Based on the local text features of each image in the first image set and a preset first similarity threshold, the first image set is clustered to obtain multiple second image clusters;

[0135] A representative image is selected from each of the plurality of second image clusters, and a second image set is obtained based on the representative image of each second image cluster;

[0136] Based on the local text features of each image in the second image set and a preset second similarity threshold, the second image set is clustered to obtain the plurality of first image clusters; wherein, the second similarity threshold is less than the first similarity threshold.

[0137] In some embodiments, the second clustering unit 620 is used for:

[0138] A representative image is selected from each of the plurality of first image clusters, and a third image set is obtained based on the representative image of each first image cluster;

[0139] Based on the global texture features of each image in the third image set, the third image set is clustered to obtain the multiple target image clusters.

[0140] In some embodiments, such as Figure 7 As shown, when the summary information is a courseware in a predetermined format, it also includes an output module 710, which is used to output the learning video and the courseware to the user device so as to associate and display the courseware and the learning video in the video recommendation page of the user device.

[0141] In this embodiment, the specific implementation methods and beneficial effects of each module or unit are as described above, and will not be repeated here.

[0142] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0143] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0144] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0145] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0146] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a method for summarizing a learning video. For example, in some embodiments, a method for summarizing a learning video can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the method for summarizing a learning video described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured, by any other suitable means (e.g., by means of firmware), to perform a method for generating a summary of a learning video.

[0147] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0148] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0149] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0150] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0151] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0152] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0153] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0154] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating summaries from learning videos, comprising: Based on the learning video, a first set of images is obtained; The learning content features of each image in the first image set are extracted; wherein, the learning content features include local text features extracted based on a feature point detection algorithm and global texture features extracted based on a convolutional neural network; the local text features include multiple text feature points, which include detailed features of the text characterized by low-level image information, including scale, orientation, and size; the global texture features include texture variation features and color features of the illustration; The first image set is clustered based on the local text features of each image to obtain multiple first image clusters. Specifically, a cluster is created for the first image in the first image set as the current cluster. Other images in the first image set are then traversed sequentially. Based on the local text features of the traversed images and the local text features of the first image in the current cluster, the similarity between the traversed image and the first image in the current cluster is determined. If the similarity is greater than or equal to a preset threshold, the traversed image is added to the current cluster. If the similarity is less than the preset threshold, a new cluster is created for the traversed image, and this new cluster is used as the current cluster. A representative image is selected from each of the plurality of first image clusters, and a third image set is obtained based on the representative image of each first image cluster; wherein, the representative image is the first image in the first image cluster; Based on the global texture features of each image in the third image set, the third image set is clustered to obtain multiple target image clusters; a representative image is selected from each of the multiple target image clusters; wherein, the cosine distance between the global texture features of two images is used as the similarity between the two images, and whether the two images belong to the same cluster is determined based on whether the similarity between the two images is greater than a preset third similarity threshold; Based on the representative image corresponding to each target image cluster, the summary information of the learning video is obtained.

2. The method according to claim 1, wherein, The first image set is clustered based on the local text features of each image to obtain multiple first image clusters, including: Based on the local text features of each image in the first image set and a preset first similarity threshold, the first image set is clustered to obtain multiple second image clusters; A representative image is selected from each of the plurality of second image clusters, and a second image set is obtained based on the representative image of each second image cluster; Based on the local text features of each image in the second image set and a preset second similarity threshold, the second image set is clustered to obtain the plurality of first image clusters; wherein, the second similarity threshold is less than the first similarity threshold.

3. The method according to claim 1, wherein, The summary information is courseware in a predetermined format; the method further includes: The learning video and the courseware are output to the user device so that the courseware and the learning video are displayed together on the video recommendation page of the user device.

4. A summary generation device for learning videos, comprising: The frame extraction module is used to obtain the first set of images based on the learning video; The image clustering module is used to extract the learning content features of each image in the first image set, and to perform clustering based on the learning content features of each image to obtain multiple target image clusters; The image selection module is used to select a representative image from each of the plurality of target image clusters; The summary determination module is used to obtain summary information of the learning video based on the representative image corresponding to each target image cluster; The image clustering module includes: A feature extraction unit is used to extract the learning content features of each image in the first image set; wherein, the learning content features include local text features extracted based on a feature point detection algorithm and global texture features extracted based on a convolutional neural network; the local text features include multiple text feature points, which include detailed features of the text characterized by low-level image information, including scale, orientation, and size; the global texture features include texture variation features and color features of the illustration; The first clustering unit is used to cluster the first image set based on the local text features of each image to obtain multiple first image clusters. Specifically, it creates a cluster for the first image in the first image set as the current cluster, sequentially traverses other images in the first image set, and determines the similarity between the traversed image and the first image in the current cluster based on the local text features of the traversed image and the local text features of the first image in the current cluster. If the similarity is greater than or equal to a preset threshold, the traversed image is added to the current cluster; if the similarity is less than the preset threshold, a new cluster is created for the traversed image, and this new cluster is used as the current cluster. The second clustering unit is used to select a representative image from each of the plurality of first image clusters, and obtain a third image set based on the representative image of each first image cluster; wherein the representative image is the first image in the first image cluster; the second clustering unit is also used to cluster the third image set based on the global texture features of each image in the third image set to obtain the plurality of target image clusters; wherein the cosine distance between the global texture features of two images is used as the similarity between the two images, and whether the two images belong to the same cluster is determined based on whether the similarity between the two images is greater than a preset third similarity threshold.

5. The apparatus according to claim 4, wherein, The first clustering unit is used for: Based on the local text features of each image in the first image set and a preset first similarity threshold, the first image set is clustered to obtain multiple second image clusters; A representative image is selected from each of the plurality of second image clusters, and a second image set is obtained based on the representative image of each second image cluster; Based on the local text features of each image in the second image set and a preset second similarity threshold, the second image set is clustered to obtain the plurality of first image clusters; wherein, the second similarity threshold is less than the first similarity threshold.

6. The apparatus according to claim 4, wherein, The summary information is courseware in a predetermined format; the device also includes: The output module is used to output the learning video and the courseware to the user device, so as to display the courseware and the learning video together on the video recommendation page of the user device.

7. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-3.

8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-3.

9. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Video abstract generation method and device, electronic equipment and storage medium

    CN110650379A

  • Video clustering method and device, storage medium and electronic equipment

    CN112131430A