A method, device, electronic device and storage medium for generating subtitles
By generating subtitles of video highlights, the problem that video highlight titles in the prior art cannot accurately reflect video content, and improve video browsing efficiency and recommendation accuracy.
Patent Information
- Application Number
- CN202110387022.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-04-12
AI Technical Summary
The titles of existing video highlights usually use the ‘summary’ title, which cannot accurately reflect the content of the video clips contained in the video highlights, making it difficult for users to quickly and accurately find the video clips of interest, and the search system cannot accurately push related videos.
By extracting reference frames from each video clip of the target video, a set of candidate videos with similarity to these reference frames is selected, and the candidate video titles of these video clips are clustered to generate subtitles of each video clip, and finally a subtitle of the target video is generated.
It improves the browsing efficiency of video clips and the accuracy of video recommendations, and reduces the time and consumption of network traffic for users to find interesting content in video highlights.
Smart Images

Figure CN113705209B_ABST
Abstract
Description
Background Art
[0002] With the popularization of mobile networks and smart terminals, the cost of short video production is getting lower and lower. The number of short videos uploaded to various new media platforms every day can reach hundreds of thousands or even millions. Among them, a large number of short videos are created through secondary creation based on original videos. Video highlights are a typical short video formed by secondary creation, which is formed by editing and splicing some popular, exciting, or original videos with the same theme.
[0003] Usually, the titles of video highlights are "summary" titles, such as "Home Cooking Recipes", "Various Funny Videos", "Stars' Wonderful Moments", etc.; however, "summary" titles can only reflect the type of video highlights, but cannot reflect the specific content included in the video highlights. This results in users being unable to quickly and accurately know which video clips are specifically included in the video highlights based on the "summary" titles. For example, when seeing the title "Home Cooking Recipes", users can only know that the video collection belongs to the food category, but cannot know which home cooking is included in the video collection; for another example, when seeing the title "Stars' Wonderful Moments", users can only know that the video collection belongs to the sports category, but cannot know which stars are included in the video collection.
[0004] Figure 1 The video highlights display interface is a manually edited "synopsis-style" title, such as Figure 1 As shown in (1), interface 11 displays the information of the video account owner, such as user name, video account owner category, video account owner level, etc. Interface 12 is the display interface of the video clips, and interface 13 displays the title of the video collection "Home Cooking Recipes #5 Home Cooking Recipes #Simple and Easy to Learn #Nutritious and Delicious" as well as the number of reposts, number of comments, playback time and other information; from the title of the video collection displayed in interface 13, it is impossible to know which home cooking video clips are included in the video collection. It is necessary to watch the video collection, and when the user searches for "Kung Pao Chicken" on the platform, even if the video collection contains a video clip of the "Kung Pao Chicken" recipe, the search system of the new media platform cannot obtain the corresponding content based on the title of the video collection, resulting in the video collection being unable to be pushed to the search user. Figure 1As shown in (2), interface 21 displays the information of the video owner, interface 22 is the display interface of the video clip, and interface 23 displays the title of the video collection "Various Sports Programs # Are There Your Idols # Olympics # National Sports" and information such as the number of reposts, number of comments, and playback time; from the title of the video collection displayed in interface 23, it is impossible to know which sports events or sports stars are included in the video collection. When a user searches for "basketball", even if the video collection contains exciting scenes of basketball games, the search system of the new media platform cannot obtain the corresponding content based on the title of the collection video, resulting in the inability to push the video collection to the search user.
[0005] Depend on Figure 1 It can be seen that in order to simplify operation, many video account owners usually add a "summary" title to the video collection that can reflect the type of each video clip. The amount of information is relatively small, and users need to perform a series of operations such as dragging, fast-forwarding, and double watching to know the playback content of each video clip. In addition, since the "summary" title cannot reflect the playback content of the corresponding video clip, when users search for target videos, the new media platform cannot accurately push the target videos that users are interested in.
[0006] Therefore, when users browse video collections, they often need to drag, fast forward, and repeatedly search for different video clips to determine whether there are clips they are interested in watching. This greatly reduces browsing efficiency and the accuracy of video recommendations, and can easily cause a waste of network traffic. Summary of the invention
[0007] The embodiments of the present application provide a subtitle generation method, device, electronic device and storage medium for improving the browsing efficiency of video clips and the accuracy of recommendations, thereby saving network traffic.
[0008] According to a first aspect of an embodiment of the present application, a subtitle generation method is provided, the method comprising:
[0009] Extract corresponding reference frames from each video clip contained in the target video;
[0010] Based on each obtained reference frame, for each corresponding video segment, a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold is screened out from the original video set;
[0011] Performing title clustering on candidate video sets corresponding to the respective video clips to obtain subtitles corresponding to the respective video clips;
[0012] Based on the subtitles corresponding to the respective video clips, the subtitle of the target video is generated.
[0013] According to a second aspect of an embodiment of the present application, a subtitle generation method is provided, the method comprising:
[0014] Extract corresponding reference frames from each video clip contained in the target video;
[0015] Based on each obtained reference frame, for each corresponding video segment, a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold is screened out from the original video set;
[0016] Title clustering is performed for each candidate video set corresponding to each of the video clips to obtain a subtitle corresponding to each of the video clips.
[0017] According to a third aspect of an embodiment of the present application, a subtitle generating device is provided, the device comprising:
[0018] A frame extraction module, used to extract corresponding reference frames from each video segment contained in the target video;
[0019] A screening module is used to screen out a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold from the original video set based on each obtained reference frame and for the corresponding video segment respectively;
[0020] The generation module is used to perform title clustering on the candidate video sets corresponding to the respective video clips to obtain the subtitles corresponding to the respective video clips; and to generate the subtitles of the target video based on the subtitles corresponding to the respective video clips.
[0021] In an optional implementation, the screening module is specifically used to:
[0022] For each reference frame, the following operations are performed respectively:
[0023] Perform frame matching on one of the reference frames with each original video included in the original video set, and determine the number of matching frames corresponding to each original video;
[0024] Based on the number of matching frames corresponding to each of the original videos, the total number of frames of each of the original videos, and the total number of frames of the video segment corresponding to the one reference frame, respectively determine the similarity between each of the original videos and the video segment corresponding to the one reference frame;
[0025] A candidate video set whose similarity with the video segment corresponding to the one reference frame reaches a similarity threshold is screened out from the original video set.
[0026] In an optional implementation, the screening module is specifically used to:
[0027] Based on a preset first operator, extract a first feature vector of the reference frame, and respectively extract a second feature vector of each original frame included in each original video;
[0028] Based on the obtained first feature vector and each second feature vector, respectively determine the first frame matching degree between the one reference frame and each original frame included in each original video;
[0029] For each of the original videos, the number of original frames whose first frame matching degree meets the first preset condition is determined as the number of matching frames corresponding to the corresponding original video.
[0030] In an optional implementation, the screening module is specifically used to:
[0031] Based on the first operator and the first set step size, frequency domain transform is performed on the one reference frame to obtain a first frequency domain value set of the one reference frame, and frequency domain transform is performed on each original frame included in each original video to obtain a second frequency domain value set corresponding to each original frame;
[0032] Determine a first frequency domain value mean corresponding to the first frequency domain value set, and respectively determine a second frequency domain value mean corresponding to each second frequency domain value set;
[0033] Based on the comparison results of each frequency domain value in the first frequency domain value set and the first frequency domain value mean, the first eigenvector of the reference frame is determined, and based on the comparison results of each frequency domain value in each second frequency domain value set and the corresponding second frequency domain value mean, the second eigenvectors corresponding to each original frame are determined respectively.
[0034] In an optional implementation, the screening module is further used to:
[0035] Based on a preset second operator, extracting a third eigenvector of the one reference frame, and respectively extracting a fourth eigenvector of each original frame included in each original video, wherein the second operator is smaller than the first operator;
[0036] Based on the obtained third eigenvector and each fourth eigenvector, respectively determine a second frame matching degree between the one reference frame and each original frame included in each original video;
[0037] In each of the original videos, the original frames whose second frame matching degree does not meet the second preset condition are deleted.
[0038] In an optional implementation, the screening module is specifically used to:
[0039] If the total number of frames of one of the original videos is less than the total number of frames of the video segment corresponding to the reference frame, the similarity between the original video and the video segment corresponding to the reference frame is positively correlated with the number of matching frames corresponding to the original video and negatively correlated with the total number of frames of the original video;
[0040] If the total number of frames of an original video among the original videos is not less than the total number of frames of the video clip corresponding to the reference frame, the similarity between the original video and the video clip corresponding to the reference frame is positively correlated with the number of matching frames corresponding to the original video, and negatively correlated with the total number of frames of the video clip corresponding to the reference frame.
[0041] In an optional implementation, the generating module is specifically used to:
[0042] For each of the video clips, perform the following operations respectively:
[0043] For a video segment among the video segments, obtaining titles of each candidate video in the corresponding candidate video set;
[0044] Perform word segmentation processing on each of the obtained titles respectively to obtain a word segmentation vector set corresponding to each of the titles;
[0045] The word vector means of each word segmentation vector set are respectively used as the title vector of the corresponding candidate video;
[0046] Title clustering is performed on the title vectors of each candidate video in the candidate video set corresponding to the one video clip to obtain a subtitle corresponding to the one video clip.
[0047] In an optional implementation, the generating module is specifically used to:
[0048] Performing title clustering on the title vectors of the candidate videos to obtain at least one candidate title category;
[0049] Determine a target title category from the at least one candidate title category based on the number of title vectors associated with each candidate title category in the at least one candidate title category;
[0050] Based on the playback volume of the candidate videos corresponding to each title vector associated with the target title category and the similarity between each of the title vectors and the title vector of the target video, the subtitle corresponding to the video clip is determined.
[0051] In an optional implementation, the frame extraction module is specifically used to:
[0052] Extracting corresponding reference frames from each video segment included in the target video according to a set target frame extraction interval, wherein the target frame extraction interval is set according to the playback duration of each video segment included in the target video; or,
[0053] Based on the target playback time of the target video and the mapping relationship between the preset playback time and the number of video clips, the target number of video clips corresponding to the target playback time is determined; based on the target playback time of the target video and the corresponding target number of video clips, a target frame extraction interval is determined, and based on the target frame extraction interval, corresponding reference frames are extracted from each video clip contained in the target video, wherein the target frame extraction interval is positively correlated with the target playback time, and negatively correlated with the number of target video clips corresponding to the target playback time.
[0054] According to a fourth aspect of an embodiment of the present application, a subtitle generating device includes:
[0055] A frame extraction module, used to extract corresponding reference frames from each video segment contained in the target video;
[0056] A screening module is used to screen out a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold from the original video set based on each obtained reference frame and for the corresponding video segment respectively;
[0057] The generating module is used to perform title clustering on the candidate video sets corresponding to the respective video clips, and obtain the subtitles corresponding to the respective video clips.
[0058] According to a fifth aspect of an embodiment of the present application, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the computer program is executed by the processor, the processor implements the subtitle generation method in the embodiment of the present application.
[0059] According to a sixth aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the subtitle generation method in the embodiment of the present application is implemented.
[0060] In the embodiment of the present application, corresponding reference frames are extracted from each video clip contained in the target video, and then based on each reference frame obtained, for each video clip, a candidate video set whose similarity with the corresponding video clip reaches a similarity threshold is screened out from the original video set, and for each candidate video set corresponding to each video clip, title clustering is performed to obtain the subtitles corresponding to each video clip, and the subtitle of the target video is generated according to the subtitles of each video clip. In this way, the subtitles of the corresponding video clips can be automatically generated for the video content of each video clip contained in the target video, and the subtitle of the target video is further generated. Since the subtitle of the target video can fully display the playback content of each video clip contained in the target video, when the user searches for the target video, the video clips of interest to the user are accurately recommended based on the search words contained in the subtitle, which improves the accuracy of the video recommendation, and after the target video is pushed to the terminal, the user does not need to perform a series of operations such as dragging, fast forwarding, and double watching, and can refer to the subtitle of the target video and accurately click on the target video that is intended to be watched, thereby effectively improving the browsing efficiency of the target video, and further, saving the consumption of network traffic. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0062] Figure 1 This is an interface diagram of a video collection in the related art;
[0063] Figure 2 A schematic diagram of an implementation environment provided for an embodiment of the present application;
[0064] Figure 3a A flow chart of a subtitle generation method provided in an embodiment of the present application;
[0065] Figure 3b A flow chart of a method for determining a candidate video set of a video segment provided in an embodiment of the present application;
[0066] Figure 3c A flow chart of a method for determining the number of matching frames of an original video provided in an embodiment of the present application;
[0067] Figure 3d A flow chart of another method for determining the number of matching frames of an original video provided in an embodiment of the present application;
[0068] Figure 3e A flow chart of a method for determining a subtitle of a video clip provided in an embodiment of the present application;
[0069] Figure 3f A flowchart of a detailed method for determining a subtitle of a video clip provided in an embodiment of the present application;
[0070] Figure 4a An interface diagram of a target video provided in an embodiment of the present application;
[0071] Figure 4b An interface diagram of a reference frame extracted according to a set target frame extraction interval provided in an embodiment of the present application;
[0072] Figure 4c An interface diagram of a reference frame extracted according to a determined target frame extraction interval provided in an embodiment of the present application;
[0073] Figure 4d A schematic diagram of frame matching provided in an embodiment of the present application;
[0074] Figure 4e A schematic diagram of a candidate video set corresponding to a video clip provided in an embodiment of the present application;
[0075] Figure 4f An interface diagram for displaying subtitles of video clips provided in an embodiment of the present application;
[0076] Figure 5 Schematic diagram of the Word2vec model provided in the embodiment of the present application;
[0077] Figure 6a A subtitle interface diagram of a target video provided in an embodiment of the present application;
[0078] Figure 6b A complete schematic diagram of the subtitle display process provided in the embodiment of the present application;
[0079] Figure 7a A functional structure diagram of a subtitle generating device for a target video provided in an embodiment of the present application;
[0080] Figure 7b A functional structure diagram of a device for generating subtitles for a video clip provided in an embodiment of the present application;
[0081] Figure 8 A structural diagram of an electronic device provided in an embodiment of the present application;
[0082] Figure 9 A hardware structure diagram of the generating device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0083] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the technical solution of the present application, rather than all of the embodiments. Based on the embodiments recorded in the application documents, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the technical solution of the present application.
[0084] It should be noted that the terms "first", "second", etc., used in the documents of this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein.
[0085] In addition, the terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.
[0086] Below, some terms in the embodiments of the present disclosure are explained to facilitate understanding by those skilled in the art.
[0087] (1) In the embodiments of the present application, the term "terminal" may include smart phones, tablet computers, wearable devices, etc.
[0088] (2) In the embodiments of the present application, the term "video highlights" refers to short videos that are created through secondary creation such as editing and re-splicing of some popular, exciting, or original videos with the same theme, including video clips that are played on various new media platforms, suitable for watching on the move and in a short leisure state, and are pushed frequently, with a playback time ranging from a few seconds to several minutes.
[0089] The themes (types) of the video clips are varied, including but not limited to skill sharing, humor, fashion trends, social hot spots, street interviews, public welfare education, advertising creativity, and commercial customization. The themes of the video clips included in the same video collection are the same.
[0090] (3) In the embodiments of the present application, the term "image perception algorithm" is a general term for a class of algorithms, including the average hash algorithm (aHash), the perceptual hash algorithm (pHash), and the difference hash algorithm (dHash). A "fingerprint" string can be generated for each image, and then the fingerprint similarity between different images can be compared.
[0091] (4) In the embodiments of the present application, the term "Word2vec" is an abbreviation of Word to Vector. The Word2vec model is a group of related models used to generate word vectors, proposed by Mikolov et al. These models are shallow and two-layer neural networks used for training to reconstruct linguistic word texts. The network is represented by words and needs to guess the input words in adjacent positions. Under the assumption of the bag-of-words model in Word2vec, the order of words is not important. After training, the Word2vec model can map each word to a word vector to represent the relationship between words.
[0092] (5) In the embodiments of the present application, the term "clustering algorithm" refers to a statistical analysis method for studying (sample or indicator) classification problems, and is also an important algorithm for data mining.
[0093] The embodiments of the present application relate to artificial intelligence (AI) and machine learning technology, and are designed based on speech processing technology (Speech Technology) and machine learning (Machine Learning, ML) in artificial intelligence.
[0094] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology mainly includes computer vision technology, speech processing technology, and machine learning / deep learning.
[0095] The following is a brief introduction to the design concept of the embodiment of the present application:
[0096] In the embodiment of the present application, corresponding reference frames are extracted from each video segment contained in the target video, and then based on each reference frame obtained, for each corresponding video segment, a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold is screened out from the original video set, and title clustering is performed for each candidate video set corresponding to each video segment to obtain the subtitle corresponding to each video segment. Since the subtitle can fully display the playback content of the corresponding video segment, when the user searches for the target video, the server of the new media platform makes accurate recommendations for the video segments that the user is interested in based on the search terms contained in the subtitle, thereby improving the accuracy of the video recommendation, and after the target video is pushed to the terminal, the user does not need to perform a series of operations such as dragging, fast forwarding, and double watching, and can refer to each subtitle and accurately click on the video segment that he intends to watch, thereby effectively improving the browsing efficiency of the video segment, and further, saving the consumption of network traffic.
[0097] It should be noted that the target video in the embodiment of the present application is a video collection, including one or more video clips.
[0098] The following describes the embodiments of the present application in conjunction with the drawings in the specification. It should be understood that the embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments in the present application and the features in the embodiments may be combined with each other if there is no conflict.
[0099] Figure 2 Schematic diagram of the implementation environment provided for the embodiment of the present application; see Figure 2 As shown, the implementation environment at least includes: a terminal 201 and a server 202.
[0100] Terminal 201 can be a device such as a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. Optionally, terminal 201 and server 202 are directly or indirectly connected via wired or wireless communication, and this application does not limit this. The user sends a video search request to the new media platform server 202 via terminal 201, and the search request carries the identifier of the target video that the user is interested in. After receiving the search request, the server 202 of each new media platform processes based on the original video set and returns the target video to terminal 201, which is displayed to the user by terminal 201, and the display content includes the subtitles of each video clip contained in the target video.
[0101] Terminal 201 generally refers to one of multiple terminals, and the embodiment of the present application is only illustrated by terminal 201. Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminals are only a few, or the above terminals are dozens or hundreds, or more. The embodiment of the present application does not limit the number and type of terminals.
[0102] Server 202 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0103] See also Figure 3a As shown, in the embodiment of the present application, the specific process of generating a subtitle is as follows:
[0104] Step S301: The server extracts corresponding reference frames from each video segment included in the target video.
[0105] In a specific implementation, the playing time of each video segment included in the target video is not less than the set time and is relatively even. Accordingly, when executing step S301, the following two methods may be used but are not limited to:
[0106] Method 1: According to the set target frame extraction interval, corresponding reference frames are extracted from each video segment included in the target video, wherein the target frame extraction interval is set according to the playback duration of each video segment included in the target video.
[0107] For example, the target video interface is as follows Figure 4a As shown, the title of the target video is "Various funny collections, I wish you happiness! #Funny Videos#Funny Collections#Sand Sculpture Videos#Humorous and Fun", the target playback time is 01:58 seconds, and it contains video clips 1, 2, and 3. Among them, the playback time of video clip 1 is 00:37 seconds, the playback time of video clip 2 is 00:40 seconds, and the playback time of video clip 3 is 00:41 seconds. The playback time of each video clip accounts for a relatively even proportion of the target playback time of the target video. Based on actual experience, the target frame extraction interval is set to 00:37 seconds to ensure that the corresponding reference frame can be extracted from each video clip. Based on the set target frame extraction interval, the corresponding reference frames are extracted from video clips 1, 2, and 3 contained in the target video.
[0108] When extracting the reference frame from the target video, the target video playback time 00:00 seconds is taken as the starting point, and the reference frame 1 is extracted from the video segment 1 according to the target frame extraction interval 00:37 seconds. The extraction result is as follows: Figure 4b As shown in (1) in the figure, reference frame 2 is extracted from video clip 2. The extraction result is shown in Figure 4b As shown in (2) in the figure, reference frame 3 is extracted from video clip 3. The extraction result is shown in Figure 4b As shown in (3) in .
[0109] It should be noted that when setting the target frame extraction interval according to the playback duration of each video segment included in the target video, it can be set according to actual conditions, and the target frame extraction interval can be the minimum value, maximum value, median, average value, etc. of the playback duration of each video segment, so as to ensure that more video segments are covered when extracting frames according to the set target frame extraction interval. In the above embodiment of the present application, the target frame extraction interval is the minimum value of the playback duration of video segments 1, 2, and 3.
[0110] Method 2: Based on the target playback time of the target video and the mapping relationship between the preset playback time and the number of video clips, determine the target number of video clips corresponding to the target playback time; based on the target playback time of the target video and the corresponding target number of video clips, determine the target frame extraction interval, and based on the target frame extraction interval, extract corresponding reference frames from each video clip contained in the target video, wherein the target frame extraction interval is positively correlated with the target playback time, and negatively correlated with the number of target video clips corresponding to the target playback time.
[0111] In specific implementation, the target video marked with the "highlights" logo can be analyzed in advance, the number of video clips contained in the target video with different playback durations can be counted, the average number of video clips in a playback duration interval can be obtained, and a dictionary can be generated. The format of the dictionary can be "playing duration interval, number of video clips", where the number of video clips in the dictionary can be the mean, median, etc., and includes the mapping relationship between the preset playback duration and the number of video clips. Optionally, the target frame extraction interval can use the following formula:
[0112] Target frame extraction interval = target playback duration / target number of video clips
[0113] Based on the determined target frame extraction interval, corresponding reference frames are extracted from each video segment included in the target video.
[0114] For example, the video collections associated with the playback duration interval [01:00, 02:00] are the first video collection, the second video collection, and the third video collection, which contain 2, 3, and 4 video segments respectively. The average number of video segments corresponding to the playback duration interval [01:00, 02:00] is (2+3+4) / 3=3. The target playback duration is 01:58 seconds. By querying the generated dictionary, it is found that the target playback duration belongs to the playback duration interval [01:00, 02:00], the number of video segments corresponding to the target video is 3, and the target frame extraction interval is seconds, based on the determined target frame extraction interval, corresponding reference frames are extracted from video segments 1, 2, and 3 contained in the target video.
[0115] When extracting the reference frame from the target video, the target video playback time 00:00 seconds is taken as the starting point, and the reference frame 1 is extracted from the video segment 1 according to the target frame extraction interval 00:39 seconds. The extraction result is as follows: Figure 4c As shown in (1) in the figure, reference frame 2 is extracted from video clip 2. The extraction result is shown in Figure 4c As shown in (2) in the figure, reference frame 3 is extracted from video clip 3. The extraction result is shown in Figure 4c As shown in (3) in .
[0116] It should be noted that in the embodiment of the present application, the target frame extraction interval can be determined by setting the coefficient of the target playback duration and the target number of video clips according to actual needs.
[0117] Step S302 : Based on each reference frame obtained, the server selects a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold from the original video set for the corresponding video segment.
[0118] In step S302, the server database stores a large number of various original videos, which can be used as the source of each video segment in the target video. For the convenience of description, they are collectively referred to as the original video set below. Since each video segment included in the target video is formed by editing and re-joining the original video, the greater the similarity between the corresponding reference frame extracted from each video segment and the video frame in the original video, the greater the probability that the video segment comes from the corresponding original video. Therefore, for each video segment included in the target video, based on each extracted reference frame, a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold is screened out from the original video set.
[0119] In the specific implementation, when executing step S302, for each reference frame, it is necessary to perform the operation of filtering out the candidate video set. The following is only an example of any reference frame (hereinafter referred to as reference frame i) among the reference frames. Figure 3b :
[0120] Step S3021 , frame matching is performed on reference frame i in each reference frame with each original video included in the original video set, and the number of matching frames corresponding to each original video is determined respectively.
[0121] For example, for reference frame 1 extracted from video clip 1 contained in the target video, it is frame matched with the first original video, the second original video, ... the Zth original video contained in the original video set. The matching results are as follows: Figure 4d As shown, the reference frame 1 matches the first original video by X1 frames, matches the second original video by Y1 frames, ... matches the Zth original video by Z1 frames.
[0122] It should be noted that the above only takes reference frame i as reference frame 1 as an example. For other reference frames, the number of matching frames with the original video is determined in the same manner, which will not be repeated here.
[0123] In an optional implementation, when executing step S3021, a perceptual hash (pHash) algorithm may be used for frame matching.
[0124] The pHash algorithm is one of the image perception algorithms. It can generate a "fingerprint" string for each image, and then compare the "fingerprints" of different images. The closer the comparison results are, the more similar the two images are. This algorithm can be used to implement the image search function in the browser. Image perception algorithms similar to the pHash algorithm include the mean hash algorithm (aHash) and the difference hash algorithm (dHash). The differences between the three algorithms are shown in Table 1.
[0125] Table 1. Differences between pHash, aHash, and dHash algorithms
[0126]
[0127]
[0128] According to Table 1, the pHash algorithm has the highest accuracy in frame matching.
[0129] The principle of pHash algorithm is as follows:
[0130] Reduce the image size, remove the details of the image, retain the basic information of the image such as structure and brightness, and discard the image differences caused by different sizes and horizontal and vertical pixel ratios. The reduced size can be set according to actual needs, such as 8*8 (pixels), 16*16 (pixels), 32*32 (pixels), etc. In the embodiment of the present application, a reference frame and an original frame are reduced to 32*32 (pixels).
[0131] Grayscale the image to simplify its color.
[0132] Based on a preset M*M size operator, a discrete cosine transform (DCT) is performed on the reduced image. In the embodiment of the present application, M is less than 32, such as M is equal to 8 or 16. Among them, DCT is a special Fourier transform that transforms the image from the pixel domain to the frequency domain, and the operator matrix represents higher and higher frequency domain coefficients from the upper left corner to the lower right corner. In order to retain the low-frequency area in the upper left corner, except for the frequency domain coefficient in the upper left corner, the other coefficients in the operator matrix are all 0 or the difference with 0 is less than the set threshold.
[0133] Calculate the DCT mean of each image after DCT transformation. For each image, perform the following steps: Starting from the upper left corner of the image, slide with a preset operator and a set step size. Each slide will get a DCT value. Based on each DCT value, determine the DCT mean of the image after DCT transformation. The DCT transformation formula is as follows:
[0134] One-dimensional DCT transform:
[0135]
[0136]
[0137] Among them, f(i) represents the picture to be transformed, i represents the number of DCT transformations, N is the number of pixels in the picture, c(u) is the compensation coefficient of DCT transformation, u represents the generalized frequency variable, u=1,2,3,…N-1, and F(u) is the coefficient after DCT transformation.
[0138] For two-dimensional images, two-dimensional DCT transformation is performed based on one-dimensional DCT transformation:
[0139]
[0140]
[0141] Wherein, v represents a generalized frequency variable, u=1,2,3,…N-1. In the embodiment of the present application, the generalized frequency variables u and v can represent the horizontal and vertical coordinates of the two-dimensional image pixel array, and c(u) and c(v) are the horizontal and vertical pixel compensation coefficients of the DCT transform, respectively.
[0142] Transforming formula 3, we get:
[0143]
[0144] It can be seen from Formula 5 that the two-dimensional DCT transform is symmetrical, so the image can be restored by inverse DCT transform.
[0145] For each picture, the DCT value after each DCT transformation is compared with the DCT mean of the corresponding picture. If the DCT value is greater than or equal to the DCT mean, it is recorded as 1, otherwise it is recorded as 0, thereby obtaining the binary array corresponding to each picture, also called the feature vector.
[0146] Image matching is performed based on the feature vectors of each image. In an embodiment of the present application, the matching degree of the two images can be determined based on the Hamming distance between the two images. The smaller the Hamming distance, the higher the matching degree of the two images.
[0147] It should be noted that, without affecting the substantive content of the present application, the embodiments of the present application do not impose restrictive requirements on the frame matching algorithm. In addition to image perception algorithms, such as the pHash algorithm, the aHash algorithm, and the dHash algorithm, matching algorithms based on local features in image classification and object recognition technology may also be adopted, such as the Scale-invariant feature transform (SIFT) algorithm, the Speeded-Up Robust Features (SURF) algorithm, the Bag of Words (BOW) algorithm, and the like.
[0148] Based on the above principle, for each extracted reference frame, the number of matching frames corresponding to the corresponding reference frame of each original video in the original video set can be determined respectively. The following takes reference frame i and original video j as an example to describe the process of determining the number of matching frames corresponding to original video j when executing step S3021. The following steps may be used but are not limited to: Figure 3c :
[0149] Step S30211: extracting a first feature vector of the reference frame i based on a preset first operator, and extracting a second feature vector of each original frame included in the original video j.
[0150] When executing step S30211, first pre-process the reference frame i and the original frames contained in the original video j, including operations such as reducing the image size and graying. Among them, the size of the image after reduction in the embodiment of the present application is 32*32 (pixels). Set the size of the first operator to 16*16, perform DCT transformation on the 32*32 (pixel) reference frame i based on the first operator, extract the first eigenvector of the reference frame i, and perform DCT transformation on each 32*32 (pixel) original frame respectively, and extract the second eigenvector of each original frame. It can be seen from the principle of the pHash algorithm that the first eigenvector and each second eigenvector are two arrays.
[0151] In a specific implementation, first, based on the first operator and the first set step size, a frequency domain transform is performed on the reference frame i to obtain a first frequency domain value set of the reference frame i, and a frequency domain transform is performed on each original frame contained in the original video j to obtain a second frequency domain value set corresponding to each original frame; then, a first frequency domain value mean corresponding to the first frequency domain value set is determined, and a second frequency domain value mean corresponding to each second frequency domain value set is determined respectively; finally, based on a comparison result of each frequency domain value in the first frequency domain value set with the first frequency domain value mean, a first eigenvector of the reference frame i is determined, and based on a comparison result of each frequency domain value in the second frequency domain value set with the corresponding second frequency domain value mean, a second eigenvector corresponding to each original frame contained in the original video j is determined respectively.
[0152] It should be noted that based on the first frequency domain value set, the median of the first frequency domain value set can also be determined, and based on the comparison result between each frequency domain value in the first frequency domain value set and the median of the first frequency domain value, the first eigenvector of the reference frame i can be determined. Similarly, the second eigenvector can also be determined.
[0153] Step S30212: Determine the first frame matching degree between the reference frame i and each original frame included in the original video j based on the obtained first feature vector and each second feature vector.
[0154] When executing step S30212, image matching is performed based on the obtained first feature vector and each second feature vector, and the first Hamming distance between the first feature vector and each second feature vector is calculated respectively. Based on each first Hamming distance, the first frame matching degree between the reference frame i and each original frame included in the original video j is determined. The smaller the first Hamming distance, the higher the first matching degree between the reference frame i and the original frame.
[0155] Step S30213: for the original video j, the number of original frames whose first frame matching degree meets the first preset condition is determined as the number of matching frames corresponding to the original video j.
[0156] When executing step S30213, the first frame matching degrees between the reference frame i and each original frame included in the original video j are compared with the first Hamming threshold Q respectively. If the first matching degree is less than the first Hamming threshold Q, it indicates that the first preset condition is met; otherwise, it indicates that the first preset condition is not met. The number of original frames whose first matching degrees with the reference frame i meet the first preset condition among the original frames included in the original video j is counted, and the counted number of frames is determined as the number of matching frames corresponding to the original video j.
[0157] In an optional implementation, in order to improve the matching degree between the reference frame i and each original frame contained in the original video j, a coarse-grained operator (also called the second operator) may be first used to extract the feature vectors of the reference frame i and each original frame, and a match is performed based on the extracted feature vectors to remove some original frames; then, a fine-grained operator (also called the first operator) is used to extract the feature vectors of the reference frame i and each remaining original frame, and a match is performed again based on the extracted feature vectors, wherein the coarse-grained operator is smaller than the fine-grained operator, thereby improving the matching degree between the reference frame i and each original frame. Therefore, before executing step S30211, the following steps may also be included, see Figure 3d :
[0158] Step S30210_1, based on a preset second operator, extracting the third eigenvector of the reference frame i, and extracting the fourth eigenvector of each original frame included in the original video j, the second operator being smaller than the first operator.
[0159] When executing step S30210_1, the reference frame i and each original frame included in the original video j are first preprocessed, and the specific description can be seen in S30211. The size of the second operator is set to 8*8, which is smaller than the first operator 16*16, and the DCT transform is performed on the reduced reference frame i based on the second operator to extract the third eigenvector of the reference frame i, and the DCT transform is performed on each reduced original frame to extract the fourth eigenvector of each original frame.
[0160] In a specific implementation, first, based on the second operator and the second set step size, a frequency domain transform is performed on the reference frame i to obtain a third frequency domain value set of the reference frame i, and a frequency domain transform is performed on each original frame contained in the original video j to obtain a fourth frequency domain value set corresponding to each original frame; then, a third frequency domain value mean corresponding to the third frequency domain value set is determined, and a fourth frequency domain value mean corresponding to each fourth frequency domain value set is determined respectively; finally, based on a comparison result of each frequency domain value in the third frequency domain value set with the third frequency domain value mean, the third eigenvector of the reference frame i is determined, and based on a comparison result of each frequency domain value in the fourth frequency domain value set with the corresponding fourth frequency domain value mean, the fourth eigenvector corresponding to each original frame contained in the original video j is determined respectively.
[0161] Step S30210_2: Based on the obtained third eigenvector and each fourth eigenvector, determine the second frame matching degree between the reference frame i and each original frame included in the original video j.
[0162] When executing step S30210_2, image matching is performed based on the obtained third eigenvector and each fourth eigenvector, and the second Hamming distance between the third eigenvector and each fourth eigenvector is calculated respectively. Based on each second Hamming distance, the second frame matching degree between the reference frame i and each original frame contained in the original video j is determined.
[0163] Step S30210_3, in the original video j, delete the original frames whose second frame matching degree does not meet the second preset condition.
[0164] When executing step S30210_3, the second frame matching degree between the reference frame i and each original frame included in the original video j is compared with the second Hamming threshold Q', if the second matching degree is less than the second Hamming threshold Q', it indicates that the second preset condition is met, otherwise it indicates that the second preset condition is not met, and the number of original frames whose second matching degree with the reference frame i meets the second preset condition among the original frames included in the original video j is counted, and the counted number of frames is determined as the matching frame number of the original video j. Among them, the second Hamming threshold Q' is greater than the first Hamming threshold Q, and based on the second frame matching degree, some original frames can be roughly screened out. Further, based on the screened original frames, steps S30211-S30213 are executed to determine the original frames in the original video j with a higher matching degree with the reference frame i, and obtain the accurate matching frame number corresponding to the original video j.
[0165] Based on the above method, the number of matching frames corresponding to each original video can be determined.
[0166] It should be noted that after executing steps S302111-S302113, it is determined that the matching degree between the reference frame i and each original frame meets the usage requirements, and steps S30210_1-S30210_3 may not be executed.
[0167] It should be noted that the embodiment of the present application measures the frame matching degree by the Chinese name distance, and the frame matching degree can also be measured by the Euclidean distance. In different scenarios, if the frame matching degree is measured by other parameters other than the distance, the larger the parameter value, the higher the frame matching degree. Therefore, in different scenarios, the first preset condition can be that the parameter value is greater than the set threshold, or it can be that the parameter value is less than the set threshold. The second preset condition is the same, and will not be repeated here.
[0168] Step S3022, based on the number of matching frames corresponding to each original video, the total number of frames of each original video, and the total number of frames of the video segment corresponding to the reference frame i, respectively determine the similarity between each original video and the video segment corresponding to the reference frame i.
[0169] In a specific implementation, when executing step S3022, any one of the original videos (hereinafter referred to as original video j, j=1, 2, ...Z) is taken as an example for explanation: if the total number of frames of the original video j is less than the total number of frames of the video segment corresponding to the reference frame i, then the similarity between the original video j and the video segment corresponding to the reference frame i is positively correlated with the number of matching frames corresponding to the original video j, and negatively correlated with the total number of frames of the original video j; if the total number of frames of the original video j is not less than the total number of frames of the video segment corresponding to the reference frame i, then the similarity between the original video j and the video segment corresponding to the reference frame i is positively correlated with the number of matching frames corresponding to the original video j, and negatively correlated with the total number of frames of the video segment corresponding to the reference frame i; optionally, the following formula may be used:
[0170] Similarity = matching frames / min (total number of frames in the video clip, total number of frames in the corresponding original video)
[0171] Because an original frame in the original video can be edited, spliced, altered, and so on for multiple times when creating the target video, the total number of frames in a video clip may be greater than the total number of frames in an original video. Therefore, when determining the similarity, it is necessary to refer to the minimum value between the total number of frames in a video clip and the total number of frames in an original video.
[0172] For example, the total number of frames of the first original video, the second original video, ... the Zth original video are sum1, sum2, ... sum3 respectively, and the total number of frames of video segment 1 corresponding to reference frame 1 is SUM1, wherein, if sum1 is less than SUM1, the similarity between the first original video and video segment 1 is P1 = X1 / sum1, if sum2 is greater than SUM1, the similarity between the second original video and video segment 1 is P2 = Y1 / SUM1, and if sum3 is equal to SUM1, the similarity between the Zth original video and video segment 1 is P3 = Z1 / SUM1; similarly, it can be determined that the similarities between video segment 2 and the first original video, the second original video, ... the Zth original video are P4 = X2 / SUM2, P5 = Y2 / SUM2, P6 = Z2 / sum2 respectively; the similarities between video segment 3 and the first original video, the second original video, ... the Zth original video are P7 = X3 / sum3, P8 = Y3 / SUM2, P9 = Z3 / sum3 respectively.
[0173] Step S3023: Filter out, from the original video set, a candidate video set whose similarity with the video segment corresponding to the reference frame i reaches a similarity threshold.
[0174] In a specific implementation, the similarity threshold is set to P. When the similarity between each original video in the original video set and the video segment corresponding to the reference frame i is greater than P, it indicates that the video segment corresponding to the reference frame i originates from the corresponding original video. The corresponding original video is used as a candidate video for the video segment corresponding to the reference frame i, where one video segment corresponds to at least one candidate video, thereby screening out a candidate video set for the video segment corresponding to the reference frame i from the original video set.
[0175] For example, if P1>P, the first original video is a candidate video for video segment 1; if P5>P, the second original video is a candidate video corresponding to video segment 2; and if P9>P, the third original video is a candidate video corresponding to video segment 2.
[0176] Taking reference frame i as reference frame 1 as an example, reference frame 1 corresponds to video segment 1. There are 3 original videos in the original video set whose similarity with video segment 1 is greater than P. Therefore, the candidate video set of video segment 1 includes candidate videos 1, 2, and 3. Figure 4e As shown in the figure, each candidate video is published by a different video account owner, and each candidate video has a title that reflects the content of the corresponding candidate video. Figure 4e The characters are outlined with a thick dotted line.
[0177] Step S303: the server performs title clustering on the candidate video sets corresponding to the respective video clips to obtain the subtitles corresponding to the respective video clips.
[0178] In step S303, each video clip is derived from a candidate video in the corresponding candidate video set. Each candidate video has its own title. Since the title of the candidate video can reflect the content of the corresponding candidate video, a subtitle of the corresponding video clip can be generated based on the title of each candidate video in the candidate video set. In this way, the user can know the content of the corresponding video clip based on the subtitle, and the new media platform can recommend videos of interest to the user based on the subtitle.
[0179] In the specific implementation, when executing step S303, for reference frame i, it is necessary to perform an operation of determining the subtitle of the corresponding video segment. The following is an example of any video segment in each video segment (hereinafter referred to as video segment i, video segment i is the video segment corresponding to reference frame i). Figure 3e :
[0180] Step S3031, for video segment i, obtain the title of each candidate video in the corresponding candidate video set.
[0181] In step S3031, the server database pre-stores the titles of the original videos. Based on the identifiers of the candidate videos in the candidate video set corresponding to the video segment i, such as the candidate video ID number, the titles of the corresponding candidate videos are obtained from the database. The title of each candidate video reflects the content of the corresponding candidate video, such as Figure 4e shown.
[0182] Taking video segment 1 corresponding to reference frame 1 as an example, the titles of candidate videos 1, 2, and 3 in the candidate video set of video segment 1 are shown in Table 2.
[0183] Table 2. Titles of candidate videos for video clip 1
[0184] Candidate video Title 1 Wonderful skiing video #Snowboarding 2 Beginner skiing video for zero foundation #Alpine skiing#Easy to learn 3 Appreciation of skiing videos #Alpine skiing#The skills are really amazing!!!
[0185] Step S3032, perform word segmentation processing on each obtained title to obtain a set of word segmentation vectors corresponding to each title.
[0186] In step S3032, an existing word segmenter may be used to segment each title, and a word segmentation algorithm may be used to segment each title. The algorithm includes but is not limited to a dictionary-based word segmentation algorithm (such as a forward maximum matching method, a reverse maximum matching method, a two-way matching word segmentation method, etc.), a statistics-based machine learning algorithm (such as a Hidden Markov Model (HMM), a Conditional Random Field Algorithm (CRF), etc.). Among them, a title can be divided into multiple word segments, each of which can be mapped to a word segmentation vector, so a title corresponds to a word segmentation vector set.
[0187] In an optional implementation, when executing step S3032, the Word2vec model can be used to obtain a set of word segmentation vectors for each title. The significance of word segmentation vectors is that they convert natural language into vectors that can be understood by computers. Compared with the bag-of-words model and the term frequency-inverse document frequency (TF-IDF) model, word segmentation vectors can capture the context and semantics of word segmentation and measure the similarity between words. They play an important role in many natural language processing fields such as text classification and sentiment analysis. The Word2vec model is introduced as follows:
[0188] The Word2vec model has three layers of neural network, namely input layer, hidden layer, and output layer. Figure 5As shown in the figure, the word "skiing" is input into the input layer, and the word is represented by a 10,000-dimensional vector containing only one 1 and the rest are 0s; a 200-dimensional feature is set to represent a word segment, then the hidden layer contains 200 neurons, and the hidden layer has no activation function, that is, the neurons are linear neurons, and the weight matrix size of the hidden layer is 10,000*300; the output layer has the same dimension as the input layer, and outputs 10,000 words including "wonderful", "video", "funny", ... "double board", and uses Softmax function for linear regression, and the sum of the probabilities of each output word is 1.
[0189] It should be noted that the parameter values in the above embodiment model are only examples and can be adjusted according to actual conditions.
[0190] In the specific implementation, in step S3032, based on the above-mentioned Word2vec model, the word segments of each title in the candidate video set of video segment i are input into the Word2vec model to obtain the 200-dimensional word vector of each word segment, thereby obtaining the word segmentation vector set corresponding to each title.
[0191] Step S3033, taking the word vector means of each word segmentation vector set obtained as the title vector of the corresponding candidate video.
[0192] In step S3033, for the word segmentation vector set corresponding to any title i in each title, the mean of each word segmentation vector in the word segmentation vector set is determined according to the dimension, and the determined mean is used as the title vector of title i.
[0193] Step S3034, performing title clustering on the title vectors of the candidate videos in the candidate video set corresponding to the video segment i, and obtaining the subtitle corresponding to the video segment i.
[0194] In step S3034, each candidate video in the candidate video set corresponding to video segment i has its own title. Since different users can add different titles when forwarding the same video, the video ID and title of the same candidate video may be different. Based on the title vectors of each candidate video, the titles are clustered according to semantic similarity to obtain the subtitle corresponding to video segment i.
[0195] Optionally, the clustering algorithm may use a k-means algorithm. The k-means algorithm is a typical unsupervised clustering algorithm, which aims to cluster input data into k clusters. The training process of the k-means algorithm is as follows:
[0196] First, randomly select k cluster centroids, denoted as μ 1 , μ 2 ,…,μk , each category has a centroid, which represents the center point of the samples belonging to the same category, where μ k ∈R n , R n Represents the set of n-dimensional real numbers.
[0197] Then, repeat the following process until convergence:
[0198] For each sample s, determine the category to which the sample s belongs. The formula is as follows:
[0199] c (s) =argmin t ||x (s) -μ t || 2 Formula 6
[0200] Among them, c (s) Indicates the category closest to sample s among k categories, with values of 1, 2, …, k, x (s) represents the feature vector of sample s, which in the embodiment of the present application refers to the title vector of the candidate video, μ t represents the centroid of category t, argmin t Represents the distance μ between sample s and category t t The minimum distance.
[0201] For each category t, recalculate the centroid of category t, the formula is as follows:
[0202]
[0203] Here, r represents the total number of samples.
[0204] Using the star cluster model to explain the k-means algorithm training process, the essence is to cluster all the stars into k star clusters. The first step is to randomly select k points in the universe (or k stars) as the centroids of the k star clusters. The second step is to calculate the distance from each star to the k centroids, and select the star cluster corresponding to the nearest centroid as c. (s) , so that after the second step, each star has its own cluster; the third step: for each cluster, determine the average value of the coordinates of all the stars in each cluster, and get the center of mass μ again t ; Repeat the second and third steps until the centroid remains unchanged or the difference between the two adjacent centroids is less than the set threshold, indicating that the k-means algorithm has reached the convergence condition and the iteration ends. The k value can be set according to the actual situation.
[0205] In a specific implementation, when executing step S3034, for the candidate video set corresponding to the video segment i, the following operations are performed:Figure 3f :
[0206] Step S30341, performing title clustering on the title vectors of each candidate video to obtain at least one candidate title category.
[0207] When executing step S30341, the title vector of each candidate video is input into the k-means model, and the k-means model clusters the title vector of each candidate video to obtain k clusters, each cluster represents a candidate title category, where k is an integer greater than or equal to 1.
[0208] Step S30342: determining a target title category from at least one candidate title category based on the number of title vectors associated with each candidate title category in at least one candidate title category.
[0209] When executing step S30342, the number of candidate title categories is k, the title vectors associated with each candidate title category have similar title semantics, and the number of title vectors associated with each candidate title category is different. In specific implementation, the candidate title category with the largest number of associated title vectors can be determined as the target title category of video segment i.
[0210] Step S30343, based on the playback volume of the candidate videos corresponding to each title vector associated with the target title category, and the similarity between each title vector and the title vector of the target video, determine the subtitle corresponding to a video clip.
[0211] When executing step S30343, from the candidate videos corresponding to the various title vectors associated with the target title category, select at least one candidate video whose playback volume is greater than the set playback threshold and whose title vectors have a similarity with the title vector of the target video within a set similarity interval. Optionally, in the embodiment of the present application, the similarity interval is [0.4-0.6], so that the title of the candidate video can provide more complementary information relative to the title of the target video to reflect the content of the video segment i. Based on the title of at least one selected candidate video, determine the score of each selected title as the subtitle of the video segment i, and determine the title of the candidate video with the highest score as the subtitle of the video segment i. The scoring formula is as follows:
[0212]
[0213] Among them, the logarithmic function log() can be used to smooth the playback volume of the candidate video, and the inverse of the cosine similarity between the title vector of the candidate video and the title vector of the target video is taken. It can be said that the lower the similarity, the higher the score of the title of the corresponding candidate video, and the more complementary information can be provided to reflect the content of video clip i.
[0214] For example, taking video segment i as video segment 1, the titles of candidate videos 1 and 2 that meet the above playback volume and similarity conditions are as follows:
[0215] Candidate video 1: Skiing #Snowboarding #Wonderful video
[0216] Candidate video 2: Skiing #Skiing #Beginner video #Zero foundation
[0217] Among them, candidate video 2 and the target video have the highest scores, so the subtitle of video clip 1 is "Skiing#Skiing#Entry-level video#Zero foundation", such as Figure 4f As shown, when playing video clip 1, the subtitle of video clip 1 is displayed below the target video title and is outlined with a thick solid line. Optionally, if the subtitle is long, the subtitle of video clip 1 can be displayed in a flowing manner when playing video clip 1. Based on the subtitle of video clip 1, the user knows that video clip 1 is specifically a video related to learning skiing, and when the user searches for "funny skiing videos", the new media platform pushes the target video corresponding to video clip 1 to the terminal based on the subtitle of video clip 1.
[0218] It should be noted that when executing step S30343, when calculating the score, in addition to considering the cosine similarity between the title vector of the candidate video and the title vector of the target video, it is also possible to consider whether the title of the candidate video contains other entity information that does not appear in the target video, such as various public names of people, places, institutions and other nouns.
[0219] Step S304: the server generates a subtitle of the target video based on the subtitles corresponding to the respective video clips.
[0220] In step S304, the subtitles of the respective video clips may be arranged according to the play order of the respective video clips to obtain the subtitle of the target video. The subtitle of the target video may be in the format of {subtitle corresponding to video clip 1 - subtitle corresponding to video clip 2 - ...}, and the subtitle of the target video is displayed when the target video is played.
[0221] For example, the target video includes video clip 1, video clip 2, and video clip 3. The subtitle of video clip 1 is "Skiing#Skiing#Beginner's video#Basic", the subtitle of video clip 2 is "Falling moment#Funny", and the subtitle of video clip 3 is "Swimming#Emoji package". Then the subtitle of the target video is "Skiing#Skiing#Beginner's video#Basic-Falling moment#Funny-Swimming#Emoji package". When the interface is displayed, Figure 6a It should be noted that, in order to match the interface size of the display device, when the subtitle of the target video is too long, the subtitle of the target video can be displayed in a flowing manner.
[0222] In step S304, the subtitle of the target video can also be generated in the form of a key-value pair based on the correspondence between each video segment and the subtitle of the video segment. The format of the subtitle of the target video can be {[identifier of video segment 1: subtitle corresponding to video segment 1], [identifier of video segment 2: subtitle corresponding to video segment 2], ...}. When the target video is played, the corresponding subtitle is displayed according to the identifier of the video segment being played.
[0223] For example, still taking the previous example, the subtitles of the target video are {[1: Skiing # Skiing # Getting Started Video # Zero Basics], [2: Falling Moment # Funny], [3: Swimming # Emoji]}. When displaying the target video, refer to Figure 4f As shown, when video clip 1 is played, the subtitle corresponding to video clip 1 is displayed, "Skiing#Skiing#Entry-level video#Zero basics".
[0224] Based on the above implementation, Figure 6b This is a complete subtitle display interface provided by the embodiment of the present application; taking the video clip 1 corresponding to the reference frame 1 as an example, Figure 6b The target video interface shown in (1) in FIG. 1 only displays the title of the target video; Figure 6b The target video shown in (2) in FIG. 1 includes an interface of a video segment 1; extracting a reference frame 1 from the video segment 1, performing frame matching between the extracted reference frame 1 and the video frames of each original video in the original video set, determining a candidate video set of the video segment 1 based on the number of matching frames, clustering each title in the candidate video set, and screening out candidate videos 2 that meet the playback volume and similarity conditions, such as Figure 6b As shown in (3) in the figure, the interface displays the title of candidate video 2. Figure 6b The (3) in the figure is outlined with a thick dotted line; based on the title of the candidate video 2, a subtitle of the video segment 1 is generated. The subtitle of the video segment 1 can reflect the detailed content of the video segment 1 and is displayed below the title of the target video. Figure 6b (4) is circled with a thick solid line. When a user searches for "funny skiing videos", the new media platform can recommend the target video corresponding to video clip 1 to the user's terminal based on the subtitle. After the terminal receives the recommended target video, if the user does not watch video clip 1, the user can know that video clip 1 is a skiing tutorial based on the subtitle of video clip 1.
[0225] It should be noted that Figure 6b (4) is only an example of a subtitle display of a video clip, and can also be displayed as Figure 6a The interface shown.
[0226] Based on the same inventive concept, the present application embodiment provides a subtitle generation device. Figure 7aAs shown, it is a structural schematic diagram of a subtitle generating device for a target video, which may include:
[0227] The frame extraction module 701 is used to extract corresponding reference frames from each video segment included in the target video;
[0228] A screening module 702 is used to screen out a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold from the original video set based on each obtained reference frame and for the corresponding video segment respectively;
[0229] The generation module 703 is used to perform title clustering on the candidate video sets corresponding to each video clip, obtain the subtitles corresponding to each video clip, and generate the subtitles of the target video based on the subtitles corresponding to each video clip.
[0230] Optionally, the screening module 702 is specifically used for:
[0231] For each reference frame, perform the following operations:
[0232] Perform frame matching on one of the reference frames with each original video included in the original video set, and determine the number of matching frames corresponding to each original video;
[0233] Based on the number of matching frames corresponding to each original video, the total number of frames of each original video, and the total number of frames of the video segment corresponding to a reference frame, respectively determine the similarity between each original video and the video segment corresponding to a reference frame;
[0234] A candidate video set whose similarity with a video segment corresponding to a reference frame reaches a similarity threshold is screened out from the original video set.
[0235] Optionally, the screening module 702 is specifically used for:
[0236] Based on a preset first operator, extract a first feature vector of a reference frame, and respectively extract a second feature vector of each original frame contained in each original video;
[0237] Based on the obtained first feature vector and each second feature vector, determining first frame matching degrees between a reference frame and each original frame included in each original video respectively;
[0238] For each original video, the number of original frames whose first frame matching degree meets the first preset condition is determined as the number of matching frames corresponding to the corresponding original video.
[0239] Optionally, the screening module 702 is specifically used for:
[0240] Based on the first operator and the first set step size, a reference frame is transformed in the frequency domain to obtain a first frequency domain value set of the reference frame, and each original frame included in each original video is transformed in the frequency domain to obtain a second frequency domain value set corresponding to each original frame;
[0241] Determine a first frequency domain value mean corresponding to the first frequency domain value set, and respectively determine a second frequency domain value mean corresponding to each second frequency domain value set;
[0242] Based on the comparison results of each frequency domain value in the first frequency domain value set and the first frequency domain value mean, a first eigenvector of a reference frame is determined, and based on the comparison results of each frequency domain value in each second frequency domain value set and the corresponding second frequency domain value mean, the second eigenvectors corresponding to each original frame are determined respectively.
[0243] Optionally, the screening module 702 is further used for:
[0244] Based on a preset second operator, extracting a third eigenvector of a reference frame, and respectively extracting fourth eigenvectors of each original frame contained in each original video, wherein the second operator is smaller than the first operator;
[0245] Based on the obtained third eigenvector and each fourth eigenvector, determining a second frame matching degree between a reference frame and each original frame included in each original video;
[0246] In each original video, the original frames whose second frame matching degree does not meet the second preset condition are deleted.
[0247] Optionally, the screening module 702 is specifically used for:
[0248] If the total number of frames of an original video among the original videos is less than the total number of frames of a video segment corresponding to a reference frame, then the similarity between an original video and a video segment corresponding to a reference frame is positively correlated with the number of matching frames corresponding to an original video, and negatively correlated with the total number of frames of an original video;
[0249] If the total number of frames of an original video among the original videos is not less than the total number of frames of a video segment corresponding to a reference frame, then the similarity between an original video and a video segment corresponding to a reference frame is positively correlated with the number of matching frames corresponding to the original video, and negatively correlated with the total number of frames of the video segment corresponding to the reference frame.
[0250] Optionally, the generating module 703 is specifically used for:
[0251] For each video clip, do the following:
[0252] For a video segment in each video segment, obtaining the title of each candidate video in the corresponding candidate video set;
[0253] Perform word segmentation processing on each obtained title respectively to obtain a set of word segmentation vectors corresponding to each title;
[0254] The word vector mean of each word vector set is taken as the title vector of the corresponding candidate video;
[0255] The title vectors of each candidate video in the candidate video set corresponding to a video clip are clustered to obtain a subtitle corresponding to the video clip.
[0256] Optionally, the generating module 703 is specifically used for:
[0257] Performing title clustering on the title vectors of each candidate video to obtain at least one candidate title category;
[0258] Determine a target title category from at least one candidate title category based on the number of title vectors associated with each candidate title category in at least one candidate title category;
[0259] Based on the playback volume of candidate videos corresponding to each title vector associated with the target title category and the similarity between each title vector and the title vector of the target video, a subtitle corresponding to a video clip is determined.
[0260] Optionally, the frame extraction module 701 is specifically used for:
[0261] According to the set target frame extraction interval, corresponding reference frames are extracted from each video segment included in the target video, wherein the target frame extraction interval is set according to the playback duration of each video segment included in the target video; or,
[0262] Based on the target playback time of the target video and the mapping relationship between the preset playback time and the number of video clips, the target number of video clips corresponding to the target playback time is determined; based on the target playback time of the target video and the corresponding number of target video clips, the target frame extraction interval is determined, and based on the target frame extraction interval, corresponding reference frames are extracted from each video clip contained in the target video, wherein the target frame extraction interval is positively correlated with the target playback time, and negatively correlated with the number of target video clips corresponding to the target playback time.
[0263] For the convenience of description, the above parts are divided into modules according to their functions and described separately. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software or hardware.
[0264] Those skilled in the art will appreciate that various aspects of the present application may be implemented as a system, method or program product. Therefore, various aspects of the present application may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to as "circuit", "module" or "system" herein.
[0265] Regarding the device in the above embodiment, the specific execution mode of each module therein has been described in detail in the embodiment of the method, and will not be elaborated here.
[0266] Based on the same inventive concept, the present application embodiment provides a subtitle generation device. Figure 7b As shown, it is a schematic diagram of the structure of a device for generating a subtitle of a video clip, which may include:
[0267] The frame extraction module 704 is used to extract corresponding reference frames from each video segment included in the target video;
[0268] A screening module 705 is used to screen out a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold from the original video set based on each obtained reference frame and for the corresponding video segment respectively;
[0269] The generating module 706 is used to perform title clustering on the candidate video sets corresponding to each video segment, and obtain the subtitle corresponding to each video segment.
[0270] For the device in the above embodiment, the specific implementation of each module is shown in Figure 7a , which will not be explained in detail here.
[0271] Figure 8 is a block diagram of an electronic device 800 according to an exemplary embodiment, the electronic device comprising:
[0272] Processor 801;
[0273] A memory 802 for storing instructions executable by the processor 801;
[0274] The processor 801 is configured to execute instructions to implement the subtitle generation method in the embodiment of the present application, for example Figures 3a to 3f Follow the steps shown in .
[0275] After introducing the method and device for generating a subtitle according to an exemplary embodiment of the present application, next, a generating device according to another exemplary embodiment of the present application is introduced.
[0276] In some possible implementations, the generating device according to the present application may include at least one processing unit and at least one storage unit. The storage unit stores program code, and when the program code is executed by the processing unit, the processing unit executes the steps of the subtitle generating method described above in the embodiment of the present application. For example, the processing unit may execute the following steps: Figures 3a to 3f Follow the steps shown in .
[0277] Refer to the following Figure 9 The generating device 900 according to this embodiment of the present application is described. Figure 9 The displayed generating device is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0278] like Figure 9 As shown, the generating device is in the form of a general computing device. The components of the generating device may include but are not limited to: at least one processing unit 901, at least one storage unit 902, and a bus 903 connecting different system components (including the storage unit 902 and the processing unit 901).
[0279] Bus 903 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a processor, or a local bus using any of a variety of bus architectures.
[0280] The storage unit 902 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922 , and may further include a read-only memory (ROM) 923 .
[0281] The storage unit 902 may also include a program / utility 925 having a set (at least one) of program modules 924, such program modules 924 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0282] The generating device may also communicate with one or more external devices 904 (e.g., keyboards, pointing devices, etc.), may also communicate with one or more devices that enable a user to interact with the generating device, and / or may communicate with any device that enables the generating device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication may be performed via an input / output (I / O) interface 905. Furthermore, the generating device may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 906. As shown, the network adapter 906 communicates with other modules for the generating device via a bus 903. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the generating device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0283] The embodiment of the present application also provides a computer storable medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above-mentioned subtitle generation method are implemented.
[0284] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0285] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0286] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the processFigure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0287] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0288] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A subtitle generation method, characterized in that: The method comprises: Extract corresponding reference frames from each video clip contained in the target video; Based on each reference frame obtained, for each corresponding video segment, a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold is selected from the original video set; For each video clip, respectively perform: performing title clustering on the title vectors of each candidate video in the candidate video set corresponding to a video clip to obtain at least one candidate title category, and determining a target title category from the at least one candidate title category based on the number of title vectors associated with each candidate title category in the at least one candidate title category, and determining a subtitle corresponding to the video clip based on the playback volume of the candidate videos corresponding to each title vector associated with the target title category and the similarity between each title vector and the title vector of the target video; Based on the subtitles corresponding to the respective video clips, the subtitle of the target video is generated.
2. The method according to claim 1, characterized in that In the process of selecting a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold from the original video set based on each obtained reference frame and for each corresponding video segment, the following operations are performed for each reference frame: Perform frame matching on one of the reference frames with each original video included in the original video set, and determine the number of matching frames corresponding to each original video; Based on the number of matching frames corresponding to each of the original videos, the total number of frames of each of the original videos, and the total number of frames of the video segment corresponding to the one reference frame, respectively determine the similarity between each of the original videos and the video segment corresponding to the one reference frame; A candidate video set whose similarity with the video segment corresponding to the one reference frame reaches a similarity threshold is screened out from the original video set.
3. The method according to claim 2, characterized in that Performing frame matching on one of the reference frames with each original video included in the original video set, and determining the number of matching frames corresponding to each original video, respectively, includes: Based on a preset first operator, extract a first feature vector of the reference frame, and respectively extract a second feature vector of each original frame included in each original video; Based on the obtained first feature vector and each second feature vector, respectively determine the first frame matching degree between the one reference frame and each original frame included in each original video; For each of the original videos, the number of original frames whose first frame matching degree meets the first preset condition is determined as the number of matching frames corresponding to the corresponding original video.
4. The method according to claim 3, characterized in that Based on a preset first operator, extracting a first feature vector of the reference frame, and respectively extracting a second feature vector of each original frame included in each original video, including: Based on the first operator and the first set step size, frequency domain transform is performed on the one reference frame to obtain a first frequency domain value set of the one reference frame, and frequency domain transform is performed on each original frame included in each original video to obtain a second frequency domain value set corresponding to each original frame; Determine a first frequency domain value mean corresponding to the first frequency domain value set, and respectively determine a second frequency domain value mean corresponding to each second frequency domain value set; Based on the comparison results of each frequency domain value in the first frequency domain value set and the first frequency domain value mean, the first eigenvector of the reference frame is determined, and based on the comparison results of each frequency domain value in each second frequency domain value set and the corresponding second frequency domain value mean, the second eigenvectors corresponding to each original frame are determined respectively.
5. The method according to claim 3, characterized in that Before extracting the first feature vector of the reference frame based on the preset first operator and respectively extracting the second feature vectors of the original frames contained in the original videos, the method further includes: Based on a preset second operator, extracting a third eigenvector of the one reference frame, and respectively extracting a fourth eigenvector of each original frame included in each original video, wherein the second operator is smaller than the first operator; Based on the obtained third eigenvector and each fourth eigenvector, respectively determine a second frame matching degree between the one reference frame and each original frame included in each original video; In each of the original videos, the original frames whose second frame matching degree does not meet the second preset condition are deleted.
6. The method according to any one of claims 2 to 5, characterized in that: Based on the number of matching frames corresponding to each of the original videos, the total number of frames of each of the original videos, and the total number of frames of the video segment corresponding to the one reference frame, respectively determining the similarity between each of the original videos and the video segment corresponding to the one reference frame, including: If the total number of frames of one of the original videos is less than the total number of frames of the video segment corresponding to the reference frame, the similarity between the original video and the video segment corresponding to the reference frame is positively correlated with the number of matching frames corresponding to the original video and negatively correlated with the total number of frames of the original video; If the total number of frames of an original video among the original videos is not less than the total number of frames of the video clip corresponding to the reference frame, the similarity between the original video and the video clip corresponding to the reference frame is positively correlated with the number of matching frames corresponding to the original video, and negatively correlated with the total number of frames of the video clip corresponding to the reference frame.
7. The method according to any one of claims 1 to 5, characterized in that: The title vector of each candidate video in the candidate video set corresponding to the video clip is obtained in the following manner: For each candidate video, execute respectively: Get the title of a candidate video; Perform word segmentation processing on the obtained title to obtain a word segmentation vector set corresponding to the title of the candidate video; The obtained word vector mean of the word segmentation vector set is used as the title vector of the candidate video.
8. The method according to claim 7, characterized in that From each video clip contained in the target video, the corresponding reference frames are extracted respectively, including: Extracting corresponding reference frames from each video segment included in the target video according to a set target frame extraction interval, wherein the target frame extraction interval is set according to the playback duration of each video segment included in the target video; or, Based on the target playback time of the target video and the mapping relationship between the preset playback time and the number of video clips, the target number of video clips corresponding to the target playback time is determined; based on the target playback time of the target video and the corresponding target number of video clips, a target frame extraction interval is determined, and based on the target frame extraction interval, corresponding reference frames are extracted from each video clip contained in the target video, wherein the target frame extraction interval is positively correlated with the target playback time, and negatively correlated with the number of target video clips corresponding to the target playback time.
9. A subtitle generation method, characterized in that: The method comprises: Extract corresponding reference frames from each video clip contained in the target video; Based on each reference frame obtained, for each corresponding video segment, a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold is selected from the original video set; For each video clip, the following steps are performed respectively: performing title clustering on the title vectors of each candidate video in the candidate video set corresponding to a video clip to obtain at least one candidate title category, and determining a target title category from the at least one candidate title category based on the number of title vectors associated with each candidate title category in the at least one candidate title category, and determining a subtitle corresponding to the video clip based on the number of play times of the candidate videos corresponding to each title vector associated with the target title category and the similarity between each title vector and the title vector of the target video.
10. A subtitle generating device, characterized in that: The device comprises: A frame extraction module, used to extract corresponding reference frames from each video segment contained in the target video; A screening module is used to screen out a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold from the original video set based on each obtained reference frame and for the corresponding video segment respectively; A generation module is used to perform, for each video clip, the following operations: performing title clustering on the title vectors of each candidate video in a candidate video set corresponding to a video clip to obtain at least one candidate title category, and determining a target title category from the at least one candidate title category based on the number of title vectors associated with each candidate title category in the at least one candidate title category, determining a subtitle corresponding to the video clip based on the number of play times of candidate videos corresponding to each title vector associated with the target title category and the similarity between each title vector and the title vector of the target video; and generating a subtitle of the target video based on the subtitles corresponding to each video clip.
11. A subtitle generating device, characterized in that: The device comprises: A frame extraction module, used to extract corresponding reference frames from each video segment contained in the target video; A screening module is used to screen out a candidate video set whose similarity with the corresponding video segment reaches a similarity threshold from the original video set based on each obtained reference frame and for the corresponding video segment respectively; A generation module is used to perform, for each video clip, the following operations: performing title clustering on the title vectors of each candidate video in a candidate video set corresponding to a video clip to obtain at least one candidate title category, and determining a target title category from the at least one candidate title category based on the number of title vectors associated with each candidate title category in the at least one candidate title category, and determining a subtitle corresponding to the video clip based on the number of play times of candidate videos corresponding to each title vector associated with the target title category and the similarity between each title vector and the title vector of the target video.
12. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the computer program is executed by the processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Systems and methods for identifying matching content
CN109643320A
Multimedia data processing method and device and storage medium
CN110598014A
Video clustering method and device, storage medium and electronic equipment
CN112131430A