Video processing method, device, electronic device, storage medium and program product
Determining the key frames of the video through similar content analysis solves the problem of low video processing efficiency, and achieving rapid and low-impact video content turning points.
Patent Information
- Application Number
- CN202210901862.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2042-07-28
AI Technical Summary
In the prior art, video processing efficiency is low, and adding advertisements to manually determine the turning point of video content affects the user's impression.
By acquiring the video collection, performing similar content analysis, determining the similar clips and positions of the target video and other videos, determining the target clip based on the number, and displaying preset information at the key frame position of the target video.
Quickly determine the turning point of video content, reduce labor consumption, improve video processing efficiency, and reduce the impact on user perception.
Smart Images

Figure CN115278300B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a video processing method, device, electronic device, storage medium, and program product. Background Art
[0002] In recent years, with the advancement of technology and the development of the times, video capture devices have become increasingly popular and video dissemination has become increasingly widespread. Video, a multimedia technology that integrates visual and auditory senses, has become an indispensable part of people's lives. Currently, before uploading videos to video players, advertisements need to be added.
[0003] The current method for adding advertisements typically involves manually processing the video. This involves manually determining the video frame at the transition point in the video. The transition point is the point where two different parts of the video connect. For example, if the video includes a title and main content, the video frame at the transition point can be the last frame of the title or the first frame of the main content. The advertisement is then added to the position of the above video frame to ensure that the advertisement does not interrupt the same content in the video and affect the user's viewing experience. However, this method of manually processing videos is inefficient and not conducive to adding advertisements to videos. Summary of the Invention
[0004] The embodiments of the present application provide a video processing method, apparatus, electronic device, storage medium, and program product, which can improve the efficiency of video processing.
[0005] The present invention provides a video processing method, including:
[0006] Obtain a video collection, where the multiple videos in the video collection include a target video and at least one other video;
[0007] Analyze the target video and other videos for similar content, and obtain similar segments in the target video and other videos, as well as the positions of the similar segments in the target video;
[0008] Determine the number of similar fragments at each position;
[0009] Based on the quantity, the target fragment is determined among all similar fragments;
[0010] Based on the target segment, a key frame is determined in the target video so as to display preset information at a position of the key frame of the target video.
[0011] The present application also provides a video processing device, including:
[0012] A first acquisition unit is configured to acquire a video set, where the multiple videos in the video set include a target video and at least one other video;
[0013] The first analysis unit is used to analyze the target video and other videos for similar content, obtain similar segments in the target video and other videos, and the positions of the similar segments in the target video;
[0014] a quantity determining unit for determining the quantity of similar fragments at each position;
[0015] a segment determination unit, configured to determine a target segment among all similar segments based on quantity;
[0016] The first target determination unit is configured to determine a key frame in a target video based on a target segment, so as to display preset information at a position of the key frame in the target video.
[0017] In some embodiments, a target video includes a target frame set, the target frame set includes multiple target frames and a frame sequence number of each target frame, and other videos include other frame sets, the other frame sets include multiple other frames. Similar content analysis is performed on the target video and other videos to obtain similar segments in the target video and other videos, as well as positions of the similar segments in the target video, including:
[0018] Calculate the similarity between the target frame and other frames;
[0019] When the similarity meets the preset conditions, the target frame is used as the similar frame of other frames, and the frame number of the target frame is used as the frame number of the similar frame;
[0020] Determine at least one similar segment from all similar frames, where the similar segment includes at least two similar frames, and the frame sequence numbers of the at least two similar frames are continuous;
[0021] According to the frame sequence number of each similar frame in the similar segment, the position of the similar segment in the target video is determined.
[0022] In some embodiments, the other frame sets further include a frame sequence number of each other frame, and determining at least one similar segment from all similar frames includes:
[0023] Determine a frame sequence difference, where the frame sequence difference is the difference between the frame sequence number of the similar frame and the frame sequence number of the corresponding other frames;
[0024] At least one similar segment is determined from all similar frames corresponding to the same frame sequence difference.
[0025] In some embodiments, the similar segments include a first merged segment, and after determining at least one similar segment from all similar frames corresponding to the same frame sequence difference value, the method further includes:
[0026] Determine a first difference, where the first difference is an absolute value of a difference between the first frame sequence difference and the second frame sequence difference, the first frame sequence difference is any one of a plurality of frame sequence differences, and the second frame sequence difference is a frame sequence difference other than the first frame sequence difference;
[0027] When the first difference is not greater than a first preset threshold, determining a second difference, where the second difference is an absolute value of a difference between a frame sequence number of a first frame in the first similar segment and a frame sequence number of a second frame in the second similar segment, the first similar segment is a similar segment corresponding to the first frame sequence difference, the second similar segment is a similar segment corresponding to the second frame sequence difference, and the first frame is adjacent to the second frame;
[0028] When the second difference is not greater than a second preset threshold, the first similar segment and the second similar segment are merged to obtain a first merged segment.
[0029] In some embodiments, the similar segments include the second merged segment, and determining the number of similar segments at each position includes:
[0030] determining, according to positions of the third similar segment and the fourth similar segment in the target video, an overlapping segment between the third similar segment and the fourth similar segment, where the third similar segment is any one of a plurality of similar segments, the fourth similar segment is a similar segment other than the third similar segment, and the plurality of similar segments include similar segments in the target video and each of the other videos;
[0031] Merging the third similar segment and the fourth similar segment according to the overlapping segments to obtain a second merged segment;
[0032] The number of similar segments at each position is determined, where the similar segments include the second merged segment and unmerged segments, and the unmerged segments are similar segments other than the third similar segment and the fourth similar segment.
[0033] In some embodiments, determining a key frame in a target video based on a target segment includes:
[0034] Determining at least one transition frame from the target video, where the transition frame includes text and a preset background;
[0035] determining a target transition frame from at least one transition frame, wherein the target transition frame is adjacent to the target segment;
[0036] Merge all intermediate frames, target transition frames, and target segments to obtain new target segments, where the intermediate frames are the frames between the target transition frames and the target segments;
[0037] Based on the new target segment, key frames are identified in the target video.
[0038] In some embodiments, determining a target segment among all the similar segments based on the number includes:
[0039] Obtaining a preset position of a preset segment in a video, where each video in the video collection includes at least a portion of the preset segment;
[0040] Based on the quantity, candidate segments are determined among all similar segments;
[0041] Comparing the position of the candidate segment in the target video with the preset position to obtain the distance between the candidate segment and the preset segment;
[0042] According to the distance, a target segment is determined from multiple candidate segments.
[0043] In some embodiments, determining a key frame in a target video based on a target segment includes:
[0044] Obtaining preset text in a target video, where the preset text is associated with a target frame in the target video;
[0045] Determining target text from preset texts, where the target text is used to indicate the playback order of the target video in the video collection;
[0046] According to the target text, a key frame is determined in the target video, where the key frame is a target frame associated with the target text.
[0047] The present application also provides a video processing method, including:
[0048] Get video and preset information;
[0049] Perform similar content analysis on two adjacent frames in the video to obtain the similarity between the two adjacent frames;
[0050] If the similarity between the two adjacent frames is lower than a third preset threshold, determining a key frame between the two adjacent frames;
[0051] Determine a plot segment in the video, where the plot segment includes all frames between two adjacent plot frames, and the plot frames include the first frame, all key frames, and the last frame in the video;
[0052] Calculate content similarity, which is the similarity between the plot segment and the preset information;
[0053] When the content similarity is greater than a fourth preset threshold, the preset information is displayed at the plot frame corresponding to the plot segment.
[0054] The present application also provides a video processing device, including:
[0055] A second acquiring unit, configured to acquire the video and preset information;
[0056] The second analysis unit is used to analyze the similarity between two adjacent frames in the video to obtain the similarity between the two adjacent frames;
[0057] a second target determination unit, configured to determine a key frame from the two adjacent frames if the similarity between the two adjacent frames is lower than a third preset threshold;
[0058] A plot determination unit is used to determine a plot segment in a video, where the plot segment includes all frames between two adjacent plot frames, and the plot frames include the first frame, all key frames, and the last frame in the video;
[0059] A similarity calculation unit, used to calculate content similarity, where the content similarity is the similarity between the plot segment and the preset information;
[0060] The display unit is configured to display the preset information at the plot frame corresponding to the plot segment when the content similarity is greater than a fourth preset threshold.
[0061] In some embodiments, if the similarity between two adjacent frames is lower than a third preset threshold, determining a key frame between the two adjacent frames includes:
[0062] Get the preset sentence corresponding to the video;
[0063] If the similarity between the two adjacent frames is lower than a third preset threshold, performing content recognition processing on the audio content corresponding to each video frame in the two adjacent frames to obtain recognition text corresponding to the two adjacent frames;
[0064] Determine, from the preset sentences, a target sentence that is identical to the recognition text corresponding to the two adjacent frames;
[0065] According to the target sentence, the key frame is determined between two adjacent frames.
[0066] In some embodiments, determining a key frame between two adjacent frames according to a target sentence includes:
[0067] When the target sentence is adjacent to the preset symbol, one of the two adjacent frames corresponding to the target sentence is used as the key frame.
[0068] In some embodiments, determining a key frame between two adjacent frames according to a target sentence includes:
[0069] When the target sentence is not adjacent to the preset symbol, content recognition processing is performed on the audio content corresponding to other video frames in the video to obtain recognized text corresponding to the other video frames, where the other video frames are the video frames after the two adjacent frames in the video;
[0070] Determining other sentences in the preset sentences that are identical to the recognized texts corresponding to the other video frames;
[0071] When other sentences are adjacent to the preset symbol, other video frames corresponding to the other sentences are used as key frames.
[0072] An embodiment of the present application also provides an electronic device, comprising a memory storing a plurality of instructions; the processor loads instructions from the memory to execute the steps of any one of the video processing methods provided in the embodiments of the present invention.
[0073] An embodiment of the present application further provides a computer-readable storage medium, which stores a plurality of instructions suitable for loading by a processor to execute the steps of any one of the video processing methods provided in the embodiments of the present invention.
[0074] An embodiment of the present application further provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of any one of the video processing methods provided in the embodiments of the present invention.
[0075] An embodiment of the present application can obtain a video collection, wherein the multiple videos in the video collection include a target video and at least one other video; perform similar content analysis on the target video and the other videos to obtain similar segments between the target video and the other videos, as well as the positions of the similar segments in the target video; determine the number of similar segments at each position; based on the number, determine the target segment among all similar segments; based on the target segment, determine the key frame in the target video, so as to display preset information at the position of the key frame of the target video.
[0076] In the present application, the positions of similar segments in different target videos may be different. Through similar content analysis, it can be known that similar segments appear in the target video and other videos together. The position of the similar segment in the target video is located, and the similar segment that appears multiple times at the same position is used as the target segment in the target video. Since the target segment appears repeatedly, it can be known that the target segment will not affect the content other than the target segment in the target video. The target segment can be used to determine the key frame in the target video. The key frame is the content turning point in the target video. By displaying preset information at the position of the key frame, the impact on the user's perception can be reduced. The video processing method of the present application can quickly determine the content turning point in the target video, and there is no need to spend manpower to determine the content turning point during the video viewing process. Therefore, the present application improves the efficiency of video processing.
[0077] The embodiment of the present application can also obtain a video and preset information; perform similar content analysis on two adjacent frames in the video to obtain the similarity between the two adjacent frames; if the similarity between the two adjacent frames is lower than a third preset threshold, determine a key frame in the two adjacent frames; determine a plot segment in the video, the plot segment includes all frames between two adjacent plot frames, and the plot frames include the first frame, all key frames and the last frame in the video; calculate content similarity, the content similarity is the similarity between the plot segment and the preset information; when the content similarity is greater than a fourth preset threshold, display the preset information at the plot frame corresponding to the plot segment.
[0078] In the present application, if the similarity between two adjacent frames is lower than the third preset threshold, the two adjacent frames correspond to two different plots in the video respectively, and the video has undergone a plot change at the position of the two adjacent frames. Any one of the two adjacent frames or the two adjacent frames is a key frame, that is, the video has a plot turning point at the position of the key frame. In this way, the video can be divided into plots by plot frames (the first frame of the video, the key frame and the last frame of the video) to obtain the plot segments corresponding to the same plot in the video, and preset information similar to the plot segment can be displayed at the plot frame corresponding to the plot segment, so that the displayed preset information is not abrupt relative to the plot segment, which can reduce the impact on the user's perception. The video processing method of the present application does not require manpower to add preset information to the video. Therefore, the present application can improve the efficiency of video processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0080] Figure 1a Schematic diagram of a video processing method according to an embodiment of the present invention;
[0081] Figure 1b Schematic diagram of a video processing method according to an embodiment of the present invention;
[0082] Figure 1c 2 is a schematic diagram of the results of the two video processing methods provided in the embodiments of the present application;
[0083] Figure 1d Schematic diagram of the video processing method provided in the embodiment of the present application;
[0084] Figure 2 Schematic diagram of the video processing method provided in the embodiment of the present application;
[0085] Figure 3a This is a schematic diagram of the structure of the model training provided in the embodiment of the present application;
[0086] Figure 3b is a schematic diagram of the model structure provided in the embodiment of the present application;
[0087] Figure 3c is a schematic diagram of the model structure provided in the embodiment of the present application;
[0088] Figure 3d is a schematic diagram of the model structure provided in the embodiment of the present application;
[0089] Figure 4a This is a structural diagram of the application of the video processing method provided in an embodiment of the present application to identify the opening and ending scenes of a video;
[0090] Figure 4b This is a structural diagram of the application of the video processing method provided in an embodiment of the present application to identify the opening and ending scenes of a video;
[0091] Figure 4c Schematic diagram of a scene of segment merging in a video processing method provided in an embodiment of the present application;
[0092] Figure 4d This is a structural diagram of a segment identifying a producer's title in a video processing method provided in an embodiment of the present application;
[0093] Figure 4e This is a structural diagram of the application of the video processing method provided in the embodiment of the present application to identify the plot segments of a video;
[0094] Figure 5 This is a schematic diagram of the first structure of the video processing device provided in an embodiment of the present application;
[0095] Figure 6 This is a schematic diagram of the second structure of the video processing device provided in an embodiment of the present application;
[0096] Figure 7 It is a structural diagram of the server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0097] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0098] Embodiments of the present application provide a video processing method, apparatus, electronic device, storage medium, and program product.
[0099] The video processing device can be integrated into an electronic device, such as a terminal or a server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, desktop computer, smart TV, smart car terminal, etc. The server can be a single server, a server cluster composed of multiple servers, or a cloud server.
[0100] In some embodiments, the video processing device may also be integrated into multiple electronic devices. For example, the video processing device may be integrated into multiple servers, and the video processing method of the present application may be implemented by the multiple servers.
[0101] In some embodiments, the server may also be implemented in the form of a terminal.
[0102] For example, reference Figure 1a The server can obtain a video collection, where the multiple videos in the video collection include a target video and at least one other video; perform similar content analysis on the target video and the other videos to obtain similar segments in the target video and the other videos, as well as the positions of the similar segments in the target video; determine the number of similar segments at each position; based on the number, determine a target segment among all similar segments; and based on the target segment, determine a key frame in the target video so that preset information is displayed at the position of the key frame of the target video. Thus, after a terminal accesses the server, the video with the preset information added at the position of the key frame can be obtained from the server and displayed.
[0103] Among them, the video collection is a TV series, a TV series includes multiple episodes, each episode includes an opening and an ending, and some videos in the TV series also include a previous episode before the opening. In order to avoid making the video too long when adding the previous episode, or shortening the duration of other content in the video except the opening and ending, part of the opening or ending will be cut off to make the length of each episode in the TV series smaller. In this way, the position and duration of the opening and ending in different videos may be different.
[0104] Through similar content analysis, the present application can know similar clips that appear in the target video and other videos at the same time, and can locate the position of the similar clip in the target video, and use the similar clips that appear multiple times in the same position as the target clip (i.e., the beginning or end of the clip) in the target video. Since the target clip appears repeatedly, it can be known that the target clip will not affect the content other than the target clip in the target video. The target clip can be used to determine the key frame in the target video. The key frame is the content turning point in the target video. By displaying preset information at the position of the key frame, the impact on the user's perception can be reduced. The video processing method of the present application can quickly determine the content turning point in the target video, and there is no need to spend manpower to determine the content turning point during the video viewing process. Therefore, the present application improves the efficiency of video processing.
[0105] For example, reference Figure 1b The server can also obtain a video and preset information; perform similar content analysis on two adjacent frames in the video to obtain a similarity between the two adjacent frames; if the similarity between the two adjacent frames is lower than a third preset threshold, determine a key frame in the two adjacent frames; determine a plot segment in the video, where the plot segment includes all frames between the two adjacent plot frames, and the plot frames include the first frame, all key frames, and the last frame in the video; calculate content similarity, which is the similarity between the plot segment and the preset information; and when the content similarity is higher than a fourth preset threshold, display the preset information at the plot frame corresponding to the plot segment. Thus, after the terminal accesses the server, it can obtain the video with the preset information added at the location of the key frame from the server and display the video.
[0106] In the present application, if the similarity between two adjacent frames is lower than a third preset threshold, the two adjacent frames correspond to two different plots in the video, and the video undergoes a plot transition at the position of the two adjacent frames. Any one of the two adjacent frames or the two adjacent frames is a key frame, that is, the video has a plot turning point (content turning point) at the position of the key frame. In this way, the video can be divided into plots by plot frames (the first frame of the video, the key frame, and the last frame of the video) to obtain the plot segments corresponding to the same plot in the video, and preset information similar to the plot segment can be displayed at the plot frame corresponding to the plot segment, so that the displayed preset information is not abrupt relative to the plot segment, which can reduce the impact on the user's perception. The video processing method of the present application does not require manpower to add preset information to the video. Therefore, the present application can improve the efficiency of video processing.
[0107] In some embodiments, reference Figure 1c , in operation Figure 1a When using the video processing method in Figure 1b Video processing methods in .
[0108] It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0109] In this embodiment, a video processing method is provided, such as Figure 1d As shown, the video processing method can be executed by an electronic device, and the specific process of the video processing method can be as follows:
[0110] 110. Obtain a video set, where the multiple videos in the video set include a target video and at least one other video.
[0111] Different videos in a video collection have the same name. For example, a video collection can be a TV series, a variety show, an animation, and so on.
[0112] Each video in a video collection has a corresponding playback order. For example, if the video collection is a TV series or anime, and it includes N episodes, and one of the videos is episode 1, then this video will be played first in the N episodes. For example, if the video collection is a variety show, and it includes N episodes, and one of the videos is episode 1, then this video will be played first in the N episodes.
[0113] The target video is the video currently being processed.
[0114] Other videos are videos in the video set that are referenced by the target video during video processing. For example, if the video set includes Episode 1, Episode 2, and so on, and the target video is Episode 1, then the other videos are at least one video from Episode 2 to Episode N.
[0115] In some embodiments, there are multiple ways to obtain the video collection. For example, the video collection can be obtained locally, from a cloud service, or from a local server, and so on.
[0116] 120. Perform similar content analysis on the target video and other videos to obtain similar segments in the target video and other videos, as well as positions of the similar segments in the target video.
[0117] Among them, similar segments are segments that exist in both the target video and other videos. For example, similar segments can be the opening, ending, and previous episode recaps, etc.
[0118] Position is used to locate similar segments in the target video. For example, if the position is frame 26-frame 150, the similar segment is the segment from frame 26 to frame 150 of the target video. If the position is second 2-second 7, the similar segment is the segment from second 2 to second 7 of the target video.
[0119] In some embodiments, considering that similar segments include multiple frames, the similarity between frames can be known by similarity. Thus, the similarity between the frames in the target video and the frames in other videos can be calculated to obtain the frames in the similar segments. The target video includes a target frame set, which includes multiple target frames and a frame sequence number of each target frame. The other videos include other frame sets, which include multiple other frames. Step 120 includes steps 121-124 (not shown in the figure):
[0120] 121. Calculate the similarity between the target frame and other frames.
[0121] 122. When the similarity meets the preset condition, the target frame is used as the similar frame of the other frames, and the frame sequence number of the target frame is used as the frame sequence number of the similar frames.
[0122] 123. Determine at least one similar segment from all similar frames, where the similar segment includes at least two similar frames, and the frame sequence numbers of the at least two similar frames are continuous.
[0123] 124. Determine the position of the similar segment in the target video according to the frame sequence number of each similar frame in the similar segment.
[0124] The target frame is any frame in the target video. Each target frame needs to be similarity calculated with all other frames in the other videos. For example, if the target video contains 12 target frames and the other videos contain 10 other frames, the similarity calculation is performed on any of the 12 target frames with the 10 other frames.
[0125] The frame number of the target frame is used to indicate the position of the target frame in the target video. For example, if the frame number of the target frame is 10, it can mean that the target frame is at the 10th frame of the target video, or it can mean that the target frame is at the 10th second of the target video, and so on.
[0126] The "other frames" are any frames in the other video. Each "other frame" needs to be similarity calculated with all target frames in the target video. For example, if the other video contains 10 other frames and the target video contains 12 target frames, the similarity calculation is performed on any of the 10 other frames with all 12 target frames.
[0127] The similarity is used to indicate the similarity between the target frame and other frames.
[0128] The preset condition is a pre-set condition used to measure the similarity between the target frame and other frames. For example, the preset condition may be that if the similarity is greater than 0.9, then the target frame is similar to the other frames. The preset condition can be set according to the actual application scenario.
[0129] A similar frame is a target frame that is similar to another frame. For example, if the similarity between the target frame and the other frame meets a preset condition, then the target frame is a similar frame to the other frame. A similar frame can correspond to one other frame, or a similar frame can correspond to multiple other frames, and so on.
[0130] The frame number of the similar frame is the frame number of the target frame that is similar to the other frame. For example, the frame number of the similar frame can refer to the i-th frame of the similar frame in the target video, or the i-th second of the similar frame in the target video, etc.
[0131] For example, the target video includes 12 target frames, and the other videos include 10 other frames. If the preset condition is that the similarity is greater than 0.9, the target frame is similar to the other frames.
[0132] If the similarity calculation between the first target frame in the target video and the 10 other frames is performed respectively, and the obtained similarities are 0.3, 0.91, 0.8, 0.8, 0.7, 0.5, 0.4, 0.5, 0.5, 0.5, then the first target frame is similar to the second other frames in the other videos, that is, one similar frame corresponds to one other frame.
[0133] If the similarity calculation between the first target frame in the target video and the 10 other frames is performed respectively, and the obtained similarities are 0.3, 0.91, 0.95, 0.8, 0.7, 0.5, 0.4, 0.5, 0.5, 0.5, then the first target frame is similar to the second other frames and the third other frames in the other videos, that is, one similar frame corresponds to multiple other frames.
[0134] In some embodiments, there are multiple ways to calculate similarity, for example, Jaccard similarity coefficient, cosine similarity, similarity calculated by distance, Pearson correlation coefficient, etc.
[0135] In some embodiments, in order to calculate the similarity between the target frame and other frames, step 121 includes steps (1) and (2) (not shown in the figure):
[0136] (1) Feature extraction is performed on the target frame and other frames respectively to obtain a first embedding vector and a second embedding vector, wherein the first embedding vector represents the image texture in the target frame and the layout of each object in the target frame, and the second embedding vector represents the image texture in the other frames and the layout of each object in the other frames;
[0137] (2) Calculate the similarity between the target frame and other frames based on the first embedding vector and the second embedding vector.
[0138] In some embodiments, although similar frames are similar to other frames, their frame numbers may be different from those of other frames. Thus, it may be difficult to determine similar segments. To facilitate obtaining similar segments, the other frame set also includes the frame number of each other frame. Step 123 includes steps I and II (not shown in the figure):
[0139] I. Determine the frame sequence difference, which is the difference between the frame sequence number of the similar frame and the frame sequence number of the corresponding other frames;
[0140] II. Determine at least one similar segment from all similar frames corresponding to the same frame sequence difference.
[0141] The frame number of the other frame is used to indicate the position of the other frame in the other video. For example, if the frame number of the other frame is 10, it may mean that the other frame is at the 10th frame of the other video, or at the 10th second of the other video, and so on.
[0142] The frame sequence difference is the difference between the frame sequence number of the similar frame and the frame sequence number of the corresponding other frames.
[0143] For example, if the frame number refers to the i-th frame in the video, that is, the frame number of the similar frame is frame 2, and the frame number of the other frame corresponding to the similar frame is frame 4, then the frame difference is 2. The frame difference specifically refers to the frame number difference between the similar frame and the other frame. If the frame number refers to the i-th second in the video, that is, the frame number of the similar frame is second 2, and the frame number of the other frame corresponding to the similar frame is fourth second, then the frame difference is also 2. In this case, the frame difference specifically refers to the time difference between the similar frame and the other frame.
[0144] For example, in [xy], x refers to the frame number of the similar frame, and y refers to the frame number of the other frame corresponding to the similar frame x. After calculating the similarity between the target frame in the target video and the other frames in the other videos, the results are [10-11], [11-12], [50-51], [51-52], [2-4], [3-5], [4-6], [6-9], and [7-10]. The frame sequence difference between [10-11] and [11-12] is 1, the frame sequence difference between [2-4], [3-5] and [4-6] is 2, and the frame sequence difference between [6-9] and [7-10] is 3. Among them, a similar segment corresponding to a frame difference of 1 consists of target frames with frame numbers 10 and 11 in the target video, that is, the similar segment is [10,11]. Another similar segment corresponding to a frame difference of 1 consists of target frames with frame numbers 50 and 51 in the target video, that is, another similar segment is [50,51]. A similar segment corresponding to a frame difference of 2 consists of target frames with frame numbers 2, 3, and 4 in the target video, that is, the similar segment is [2,3,4]. A similar segment corresponding to a frame difference of 3 consists of target frames with frame numbers 6 and 7 in the target video, that is, the similar segment is [6,7].
[0145] In some embodiments, in order to increase the length of similar segments and reduce the number of similar segments, the similar segments include the first merged segment, and after step II in step 123, further comprising:
[0146] Determine a first difference, where the first difference is a difference between a first frame sequence difference and a second frame sequence difference, the first frame sequence difference is any one of a plurality of frame sequence differences, and the second frame sequence difference is a frame sequence difference other than the first frame sequence difference;
[0147] When the first difference is not greater than a first preset threshold, determining a second difference, where the second difference is an absolute value of a difference between a frame sequence number of a first frame in the first similar segment and a frame sequence number of a second frame in the second similar segment, the first similar segment is a similar segment corresponding to the first frame sequence difference, the second similar segment is a similar segment corresponding to the second frame sequence difference, and the first frame is adjacent to the second frame;
[0148] When the second difference is not greater than a second preset threshold, the first similar segment and the second similar segment are merged to obtain a first merged segment.
[0149] The first frame sequence difference value is any one of a plurality of frame sequence difference values. For example, if the plurality of frame sequence difference values include 1, 2, and 3, then the first frame sequence difference value is any one of 1, 2, and 3.
[0150] The second frame sequence difference is a frame sequence difference other than the first frame sequence difference. For example, if the first frame sequence difference is 1, the second frame sequence difference can be 2 or 3.
[0151] The first difference is the absolute value of the difference between the first frame sequence difference and the second frame sequence difference. For example, if the first frame sequence difference is 1 and the second frame sequence difference is 2, then the first difference is equal to 1, and so on.
[0152] The first preset threshold is used to measure the first difference, and the first preset threshold is determined based on the actual application scenario. For example, the first preset threshold may be 1. When the first difference is 1, the first difference is not greater than the first preset threshold, thereby preliminarily determining that the similar segment corresponding to the first frame sequence difference can be merged with the similar segment corresponding to the second frame sequence difference.
[0153] The second difference is the absolute value of the difference between the frame number of the first frame in the first similar segment and the frame number of the second frame in the second similar segment, where the first frame and the second frame are adjacent. For example, when the first similar segment is [10, 11] and the second similar segment is [2, 3, 4], since the first frame and the second frame are adjacent, the frame number of the first frame is 11 and the frame number of the second frame is 2, and the second difference is 9. When the first similar segment is [2, 3, 4] and the second similar segment is [6, 7], since the first frame and the second frame are adjacent, the frame number of the first frame is 4 and the frame number of the second frame is 6, and the second difference is 2.
[0154] The second preset threshold is used to measure the second difference, and the second preset threshold is determined based on the actual application scenario. For example, the second preset threshold may be 3. When the second difference is 9, [10,11] and [2,3,4] are two independent similar segments. When the second difference is 2, the second difference is no greater than the second preset threshold. Therefore, it can be determined that the similar segment corresponding to the first frame sequence difference can be merged with the similar segment corresponding to the second frame sequence difference, that is, [2,3,4] and [6,7] are merged into one similar segment.
[0155] The first merged segment is a similar segment obtained by merging the first similar segment and the second similar segment. For example, merging [2,3,4] and [6,7] into one similar segment yields [2,3,4,5,6,7].
[0156] 130. Determine the number of similar segments at each position.
[0157] The number is used to indicate the number of similar segments at the same position in the target video.
[0158] For example, a target video is analyzed for similar content with other video 1, other video 2, and other video 3 respectively, and similar segments 1a and 1b are obtained between the target video and other video 1, similar segment 2 is obtained between the target video and other video 2, and similar segment 3 is obtained between the target video and other video 3. The position of similar segment 1a in the target video is from the 26th frame to the 150th frame, the position of similar segment 1b in the target video is from the 250th frame to the 275th frame, the position of similar segment 2 in the target video is from the 26th frame to the 150th frame, and the position of similar segment 3 in the target video is from the 26th frame to the 150th frame. Then, the number of similar segments positioned from the 26th frame to the 150th frame is 3, and the number of similar segments positioned from the 250th frame to the 275th frame is 1.
[0159] In some embodiments, considering that a target frame in a target video is similar to multiple other frames in other videos, or that the target video overlaps with similar segments in different other videos, in order to reduce repeated similar segments, the similar segments include a second merged segment, and determining the number of similar segments at each position includes:
[0160] determining, according to positions of the third similar segment and the fourth similar segment in the target video, an overlapping segment between the third similar segment and the fourth similar segment, where the third similar segment is any one of a plurality of similar segments, the fourth similar segment is a similar segment other than the third similar segment, and the plurality of similar segments include similar segments in the target video and each of the other videos;
[0161] Merging the third similar segment and the fourth similar segment according to the overlapping segments to obtain a second merged segment;
[0162] The number of similar segments at each position is determined, where the similar segments include the second merged segment and unmerged segments, and the unmerged segments are similar segments other than the third similar segment and the fourth similar segment.
[0163] The third similar segment is any one of the multiple similar segments, and the multiple similar segments include similar segments in the target video and each of the other videos. For example, if the multiple similar segments include [2, 3, 4, 5, 6, 7], [3, 4, 5], [10, 11], and [4, 5, 6, 7, 8, 9], then the third similar segment is any one of [2, 3, 4, 5, 6, 7], [3, 4, 5], [10, 11], and [4, 5, 6, 7, 8, 9].
[0164] The fourth similar segment is a similar segment other than the third similar segment. For example, if the third similar segment is [2, 3, 4, 5, 6, 7], then the fourth similar segment is any one of [3, 4, 5], [10, 11], and [4, 5, 6, 7, 8, 9].
[0165] Overlapping segments are segments where the third and fourth similar segments overlap. For example, if the third similar segment is [2,3,4,5,6,7] and the fourth similar segment is [3,4,5], the overlapping segment is [3,4,5]. If the third similar segment is [2,3,4,5,6,7] and the fourth similar segment is [4,5,6,7,8,9], the overlapping segment is [4,5,6,7]. If the third similar segment is [2,3,4,5,6,7] and the fourth similar segment is [10,11], there is no overlapping segment.
[0166] The second merged segment is the segment resulting from merging the third similar segment with the overlapping segments in the fourth similar segment. For example, if the third similar segment is [2,3,4,5,6,7], the fourth similar segment is [3,4,5], and the overlapping segment is [3,4,5], the second merged segment is [2,3,4,5,6,7]. If the third similar segment is [2,3,4,5,6,7], the fourth similar segment is [4,5,6,7,8,9], and the overlapping segment is [4,5,6,7], the second merged segment is [2,3,4,5,6,7,8,9].
[0167] Unmerged segments are similar segments that do not overlap with other similar segments. For example, if the segment [10,11] in the similar segments [2,3,4,5,6,7], [3,4,5], [10,11], and [4,5,6,7,8,9] does not overlap with other similar segments, then [10,11] is an unmerged segment.
[0168] When a target frame in the target video is similar to multiple other frames in other videos at the same time, for example, the third similar segment is [2,3,4,5,6,7], the frame sequence difference corresponding to the third similar segment is 2, and the similar frame with frame number 3 in the third similar segment corresponds to the other frame with frame number 5. The fourth similar segment is [3,4,5], the frame sequence difference corresponding to the fourth similar segment is 3, and the similar frame with frame number 3 in the fourth similar segment corresponds to the other frame with frame number 6. In this way, the similar frame with frame number 3 corresponds to different other frames in different frame sequence differences, and there is a repeated segment [3,4,5] in the similar segments a and the similar segments b. In order to reduce the repeated similar segments, the first similar segment [2,3,4,5,6,7] and the second similar segment [3,4,5] are merged to obtain the second merged segment [2,3,4,5,6,7].
[0169] Or the target video overlaps with similar segments in different other videos. For example, similar segments between the target video and other video 1 include [2, 3, 4, 5, 6, 7] and [10, 11], similar segments between the target video and other video 2 include [2, 3, 4, 5, 6, 7], and similar segments between the target video and other video 3 include [2, 3, 4, 5, 6, 7]. Then, similar segments [2, 3, 4, 5, 6, 7] between the target video and other video 1, similar segments [2, 3, 4, 5, 6, 7] between the target video and other video 2, and similar segments [2, 3, 4, 5, 6, 7] between the target video and other video 3 are merged to obtain the second merged segment [2, 3, 4, 5, 6, 7].
[0170] In some embodiments, in order to accurately merge the third similar segment and the fourth similar segment, the third similar segment and the fourth similar segment are merged according to the overlapping segments to obtain a second merged segment, including:
[0171] Get the length of the overlapping segment and the length of the third similar segment;
[0172] determining a length ratio, the length ratio being a ratio of a length of the overlapping segment to a length of a third similar segment;
[0173] When the length ratio is greater than a preset target threshold, the third similar segment and the fourth similar segment are merged to obtain a second merged segment.
[0174] The length of the overlapping segment is the length corresponding to the position of the overlapping segment in the target video. For example, if the overlapping segments are [3, 4, 5], the length of the overlapping segment is 3.
[0175] The length of the third similar segment is the length corresponding to the position of the third similar segment in the target video. For example, if the third similar segment is [2, 3, 4, 5, 6, 7], the length of the overlapping segment is 5.
[0176] The length ratio is the ratio of the length of the overlapping segment to the length of the third similar segment. For example, if the length of the overlapping segment is 3 and the length of the third similar segment is 5, the length ratio is 0.6.
[0177] The preset target threshold is used to measure the length ratio, wherein the preset target threshold can be determined according to an actual application scenario.
[0178] For example, the preset target threshold is 0.5 and the length ratio is 0.6, and the length ratio is greater than the preset target threshold, and the third similar segment and the fourth similar segment are merged.
[0179] The method of merging the third similar segment and the fourth similar segment includes:
[0180] (1) When the third similar segment contains the fourth similar segment, and the length of the fourth similar segment is greater than the preset target threshold multiplied by the length of the third similar segment, the fourth similar segment is deleted and the third similar segment is retained.
[0181] (2) When the third similar segment intersects the fourth similar segment, the length of the overlapping segment is greater than a preset target threshold multiplied by the length of the third similar segment, and the number of similar frames in the fourth similar segment is greater than a preset number, the third similar segment and the fourth similar segment are merged.
[0182] (3) When the third similar segment intersects with the fourth similar segment, the length of the overlapping segment is greater than the preset target threshold multiplied by the length of the third similar segment, and the number of similar frames in the fourth similar segment is less than the preset number, the fourth similar segment is deleted and the third similar segment is retained.
[0183] (4) When the third similar segment intersects with the fourth similar segment, and the length of the overlapping segment is less than the preset target threshold multiplied by the length of the third similar segment, the fourth similar segment is deleted and the third similar segment is retained.
[0184] 140. Based on the quantity, determine the target segment among all similar segments.
[0185] The target segment is a similar segment that appears multiple times at the same location. For example, if the number of similar segments between frames 26 and 150 is 3, and the number of similar segments between frames 250 and 275 is 1, then the target segment is the similar segment between frames 26 and 150.
[0186] In some embodiments, considering that the titles related to the producer in a video are relatively fixed, in order to identify the titles related to the producer from multiple similar clips, step 140 includes:
[0187] Obtaining a preset position of a preset segment in a video, where each video in the video collection includes at least a portion of the preset segment;
[0188] determining a candidate segment among all similar segments based on the number;
[0189] Comparing the position of the target segment in the target video with the preset position to obtain the distance between the candidate segment and the preset segment;
[0190] Based on the distance, the target segment is determined in the target video.
[0191] Among them, the preset clips are clips that are pre-set to appear repeatedly in multiple videos, such as clips related to the producer, or clips related to the curtain call, etc.
[0192] The preset position is the position of the preset segment in the video when the preset segment is not deleted.
[0193] Candidate segments are similar segments that meet a preset number of times. For example, if all similar segments include similar segment A, which appears three times, similar segment B, which appears two times, and similar segment C, which appears once, then the candidate segments can be similar segment A, similar segment A and similar segment B, and so on.
[0194] The distance is the difference between the frame number of the first frame in the candidate segment and the frame number of the first frame in the preset segment, and can also be the difference between the frame number of the last frame in the candidate segment and the frame number of the last frame in the preset segment.
[0195] For example, the preset segment is [1,2,3,4,5,6,7,8,9], the candidate segment A is [2,3,4,5,6,7], and the candidate segment B is [7,8,9,10]. The distance between the preset segment [1,2,3,4,5,6,7,8,9] and the candidate segment A [2,3,4,5,6,7] is 1, and the overlapping segment between the preset segment [1,2,3,4,5,6,7,8,9] and the candidate segment B [7,8,9,10] is 6. In this way, the candidate segment A is closer to the preset segment than the candidate segment, so the candidate segment A is the target segment.
[0196] 150. Based on the target segment, determine a key frame in the target video, so as to display preset information at a position of the key frame of the target video.
[0197] The key frame is a video frame obtained from the target segment in the target video. For example, the key frame is connected to the target segment in the target video.
[0198] Preset information is information that is pre-set and displayed at the location of the key frame. For example, the preset information can be an advertisement, a progress content description, a video supplementary description, etc. The progress content description is used to explain the content of the progress bar corresponding to the target segment. For example, the progress content description can be used to indicate whether the target segment is the opening or ending credits. The video supplementary description can be used to explain a scene in the video. For example, if the scene is "Tang Dynasty Street Market", the video supplementary description is used to explain the content related to "Tang Dynasty Street Market".
[0199] Specifically, the progress content description may be displayed above or below the progress bar corresponding to the target segment.
[0200] In some embodiments, when the preset information is a supplementary description of an advertisement or video, the supplementary description of the advertisement or video is added before or after the target segment, and the advertisement is connected to the target segment.
[0201] In some embodiments, when the preset information is a progress content description, the progress content description is displayed above or below the progress bar corresponding to the target segment, wherein the progress bar corresponding to the target segment is between the frame number of the first frame and the frame number of the last frame of the target segment.
[0202] In some embodiments, considering that the video's opening credits are followed by the closing credits, and the closing credits do not affect other content of the video except the opening credits and the closing credits, the opening credits and all frames between the opening credits and the closing credits may be merged to locate the target frame. Step 150 includes:
[0203] Determining at least one transition frame from the target video, where the transition frame includes text and a preset background;
[0204] determining a target transition frame from at least one transition frame, wherein the target transition frame is adjacent to the target segment;
[0205] Merge all intermediate frames, target transition frames, and target segments to obtain new target segments, where the intermediate frames are the frames between the target transition frames and the target segments;
[0206] Based on the new target segment, key frames are identified in the target video.
[0207] A transition frame is a frame in a video that includes text and a preset background. The preset background can be determined based on the actual application scenario. For example, if the preset background is black, a frame in a video that only has a black background and text is considered a transition frame.
[0208] The target transition frame is the transition frame closest to the target segment in the target video, that is, the target transition frame is adjacent to the target segment. For example, the target transition frame is the frame corresponding to the announcement in the target video.
[0209] The intermediate frames are the frames between the target transition frame and the target segment. For example, if the target segment is located at [2, 3, 4, 5, 6, 7] in the target video and the target transition frame is numbered 10, then the intermediate frames are the frames numbered 8 and 9 in the target video.
[0210] The new target clip includes all intermediate frames, target transition frames, and the target clip.
[0211] For example, after obtaining a new target segment, the focus frame can be placed before or after the new target segment, so that the advertisement can be added before or after the new target segment, and the advertisement is connected to the new target segment.
[0212] In some embodiments, the video comprises a complete preset segment.
[0213] In some embodiments, in order to reduce the impact of the header on the length of the video, some preset segments can be retained in the video, and the video includes some preset segments.
[0214] In some embodiments, considering that a video has an announcement point, that is, when a video is displayed, the episode or episode number will be displayed, adding advertisements before or after the announcement point does not affect the main plot content of the video. Step 150 includes:
[0215] Obtaining preset text in a target video, where the preset text is associated with a target frame in the target video;
[0216] Determining target text from preset texts, where the target text is used to indicate the playback order of the target video in the video collection;
[0217] According to the target text, a key frame is determined in the target video, where the key frame is a target frame associated with the target text.
[0218] The preset text is text associated with the target video, and the preset text is associated with the target frame in the target video. For example, the preset text includes all lines in the target video, or includes all subtitles in the target video, etc.
[0219] The target text indicates the order in which the target video should be played within the video collection. For example, the target text could be episode number, episode number, etc. For example, if the target frame associated with the target text is at the 10th second in the target video, the focus frame is the frame corresponding to the 10th second.
[0220] As can be seen from the above, the embodiments of the present application can obtain a video collection, wherein the multiple videos in the video collection include a target video and at least one other video; perform similar content analysis on the target video and the other videos to obtain similar segments in the target video and the other videos, as well as the positions of the similar segments in the target video; determine the number of similar segments at each position; based on the number, determine the target segment among all similar segments; based on the target segment, determine the key frame in the target video, so as to display preset information at the position of the key frame of the target video.
[0221] Therefore, this solution can use similar segments that appear multiple times in the same position as the target segment in the target video. Since the target segment appears repeatedly, it can be seen that the target segment will not affect the content other than the target segment in the target video. The target segment can be used to determine the key frame in the target video. The key frame is the turning point of the content in the target video. By displaying preset information at the position of the key frame, the impact on the user's perception can be reduced. The video processing method of the present application can quickly determine the content turning point in the target video, and there is no need to consume manpower to determine the content turning point during the video viewing process. Therefore, the present application improves the efficiency of video processing.
[0222] In this embodiment, a video processing method is provided, such as Figure 2 As shown, the video processing method can be executed by an electronic device, and the specific process of the video processing method can be as follows:
[0223] 210. Obtain video and preset information.
[0224] A video is any video in a video collection, such as a TV series, a variety show, or a movie.
[0225] Preset information refers to information that is pre-set and ready for display. For example, the preset information can be an advertisement, a progress description, or a supplementary description of the video. The progress description is used to explain the content of the progress bar corresponding to the target segment. For example, the progress description can be used to indicate whether the target segment is the opening or ending credits. The supplementary description of the video can be used to explain a scene in the video. For example, if the scene is a "Tang Dynasty street market," the supplementary description of the video is used to explain the content related to the "Tang Dynasty street market."
[0226] 220. Perform similar content analysis on two adjacent frames in the video to obtain the similarity between the two adjacent frames.
[0227] The similarity is used to indicate the similarity between two adjacent frames.
[0228] In some embodiments, in order to obtain the similarity between two adjacent frames, step 220 includes:
[0229] Perform feature extraction on each of two adjacent frames to obtain a third embedding vector and a fourth embedding vector. The third embedding vector represents the semantics and layout of each object in the first frame of the two adjacent frames, and the fourth embedding vector represents the semantics and layout of each object in the second frame of the two adjacent frames.
[0230] The similarity between two adjacent frames in the video is calculated based on the third embedding vector and the fourth embedding vector.
[0231] 230. If the similarity between the two adjacent frames is lower than a third preset threshold, determine a key frame between the two adjacent frames.
[0232] The third preset threshold is used to measure the similarity between two adjacent frames, and the third preset threshold can be determined according to actual application scenarios. For example, the third preset threshold can be 0.45, 0.4, 0.3, and so on.
[0233] For example, when the third preset threshold is 0.5 and the similarity between two adjacent frames is 0.35, a key frame can be determined from the two adjacent frames. The key frame can be any one of the two adjacent frames or two adjacent frames.
[0234] In some embodiments, considering that a line may correspond to two dissimilar adjacent frames, in order to avoid dividing a line into two different plots when determining a key frame (a turning point in the video content), if the similarity between the two adjacent frames is lower than a third preset threshold, determining the key frame from the two adjacent frames includes:
[0235] Get the preset sentence corresponding to the video;
[0236] If the similarity between the two adjacent frames is lower than a third preset threshold, performing content recognition processing on the audio content corresponding to the two adjacent frames to obtain recognition text corresponding to the two adjacent frames;
[0237] Determine, from the preset sentences, a target sentence that is identical to the recognition text corresponding to the two adjacent frames;
[0238] According to the target sentence, the key frame is determined between two adjacent frames.
[0239] Among them, the preset sentence can include text corresponding to the audio of the video. For example, the preset sentence can be a pre-set video dialogue script, or it can be subtitles associated with the audio in the video except for the opening and ending credits. The audio includes the audio of each character in the video and the audio of the narration, etc.
[0240] The recognized text corresponding to two adjacent frames is the same as the audio content corresponding to two adjacent frames. For example, the recognized text between two adjacent frames can be subtitles in the video frames, or can be the text after audio recognition of the audio corresponding to two adjacent frames, and so on.
[0241] The target sentence is the sentence in the preset sentences that matches the recognized text corresponding to two adjacent frames. For example, the preset sentences corresponding to a video may include the first line, the second line, and so on. If the second line matches the recognized text corresponding to one of the two adjacent frames, then the second line is the target sentence.
[0242] For example, when one of the two adjacent frames contains a whole line of a preset sentence, the one of the two adjacent frames is used as the key frame; or when one of the two adjacent frames contains the last sentence of a whole line, the one of the two adjacent frames is used as the key frame, and so on.
[0243] In some embodiments, in order to identify the last sentence of a speech corresponding to a video frame, a key frame is determined between two adjacent frames according to the target sentence, including:
[0244] When the target sentence is adjacent to the preset symbol, one of the two adjacent frames corresponding to the target sentence is used as the key frame.
[0245] The preset symbol is a pre-set punctuation mark, which is used to indicate the end of a sentence in the preset sentence. For example, the preset symbol can be a period, a question mark, an exclamation mark, etc.
[0246] In some embodiments, the preset symbol is determined according to the actual application scenario.
[0247] For example, the preset sentences corresponding to the video may include the first line, the second line, and the Nth line. The second line includes sentences 1 and 2, the second line includes sentences 3 and 4, and the Nth line includes sentence n. When the target sentence is sentence 3 in the second line, the character "," adjacent to the target sentence in the preset sentence is a comma, which is not a preset symbol. Therefore, one of the two adjacent video frames is not a key frame. When the target sentence is sentence 4 in the second line, the character "." adjacent to the target sentence in the preset sentence is a period, which is a preset symbol. Therefore, one of the two adjacent video frames is a key frame.
[0248] In some embodiments, considering that the lines corresponding to two adjacent frames are in a complete line, in order to avoid interrupting the complete line when the key frame is any of the two adjacent frames, the key frame is determined in the two adjacent frames according to the target sentence, including:
[0249] When the target sentence is not adjacent to the preset symbol, content recognition processing is performed on the audio content corresponding to other video frames in the video to obtain recognized text corresponding to the other video frames, where the other video frames are the video frames after the two adjacent frames in the video;
[0250] Determining other sentences in the preset sentences that are identical to the recognized texts corresponding to the other video frames;
[0251] When other sentences are adjacent to the preset symbol, other video frames corresponding to the other sentences are used as key frames.
[0252] The other video frames are video frames following two adjacent frames in the video. For example, the other video frames are the first video frame, the second video frame, ... the Nth video frame, and so on, following two adjacent frames in the video.
[0253] The recognized texts corresponding to other video frames are the same as the audio contents corresponding to other video frames. For example, the recognized texts of other video frames may be subtitles in other video frames, or texts obtained after audio recognition of the audio corresponding to other video frames, and so on.
[0254] Other sentences are sentences in the preset sentences that have the same recognized text as other video frames. For example, the preset sentences corresponding to a video may include the first line, the second line, and so on. If the second line is the same as the recognized text corresponding to other video frames, then the second line is the target sentence.
[0255] For example, when other video frames contain a whole sentence of a preset sentence, the other video frames are used as key frames; or when other video frames contain the last sentence of a whole sentence, the other video frames are used as key frames, and so on.
[0256] For example, the preset sentences corresponding to the video may include the first line, the second line, and the Nth line. The second line includes sentences 1 and 2, the second line includes sentences 3 and 4, and the Nth line includes sentence n. When the other sentence is sentence 3 in the second line, the adjacent character in the preset sentence is a comma, which is not a preset symbol. Therefore, the video frame is not a key frame. When the other sentence is sentence 4 in the second line, the adjacent character in the preset sentence is a period, which is a preset symbol. Therefore, the video frame is a key frame.
[0257] 240. Determine a plot segment in the video, where the plot segment includes all frames between two adjacent plot frames, and the plot frames include the first frame, all key frames, and the last frame in the video.
[0258] The similarity between two adjacent frames in the plot segment is greater than a third preset threshold.
[0259] The plot frames include the first frame, all key frames and the last frame in the video arranged in the video frame playback order, the first frame is the first video frame of the video, and the last frame is the last video frame of the video.
[0260] For example, the plot frames of a video include {first frame, first key frame, second key frame, last frame}, then the plot segments of the video include plot segment 1, plot segment 2, and plot segment 3. Plot segment 1 includes all frames between the first frame and the first key frame, plot segment 2 includes all frames between the first key frame and the second key frame, and plot segment 3 includes all frames between the second key frame and the last plot segment.
[0261] 250. Calculate content similarity, where the content similarity is the similarity between the plot segment and the preset information.
[0262] The content similarity is used to indicate the similarity between the preset information and the plot segment.
[0263] In some embodiments, in order to calculate the similarity between the plot segment and the preset information, step 250 includes steps 251 to 253 (not shown in the figure):
[0264] 251. Extract features from the preset information to obtain a first feature;
[0265] 252. Extract features from the plot segment to obtain a second feature;
[0266] 253. Calculate content similarity based on the first feature and the second feature.
[0267] The first feature is used to represent the preset information. For example, the first feature can be a vector representing the preset information, or a matrix representing the preset information, etc.
[0268] The second feature is used to represent the plot segment. For example, the first feature can be a vector representing the plot segment, or a matrix representing the plot segment, etc.
[0269] In some embodiments, there are multiple ways to calculate content similarity using the first feature and the second feature, for example, Jaccard similarity coefficient, cosine similarity, similarity calculated by distance, Pearson correlation coefficient, and so on.
[0270] 260. When the content similarity is greater than a fourth preset threshold, display the preset information at the plot frame corresponding to the plot segment.
[0271] The fourth preset threshold is used to measure the content similarity to determine whether the preset segment is similar to the plot segment. The fourth preset threshold is determined according to an actual application scenario.
[0272] For example, the preset information is an advertisement. In order to make the advertisement appear unobtrusive at the location of the key frame, feature similarity can be used to determine whether the advertisement is similar to a plot segment in the video. If the advertisement is similar to the plot segment, the advertisement will be displayed at the plot frame corresponding to the plot segment.
[0273] For example, the plot segments of the video include plot segment 1, plot segment 2, and plot segment 3. Plot segment 1 includes all frames between the first frame and the first key frame, plot segment 2 includes all frames between the first key frame and the second key frame, and plot segment 3 includes all frames between the second key frame and the last plot segment. Among them, the preset information can be added after the first key frame corresponding to plot segment 1, that is, the preset information can be displayed at the plot frame corresponding to plot segment 1. The preset information can also be added before the first key frame corresponding to plot segment 2, or after the second key frame, so that the preset information can be displayed at the plot frame corresponding to plot segment 2. The preset information can also be added before the second key frame corresponding to plot segment 3, so that the preset information can be displayed at the plot frame corresponding to plot segment 3.
[0274] As can be seen from the above, the embodiment of the present application can obtain a video and preset information; perform similar content analysis on two adjacent frames in the video to obtain the similarity of the two adjacent frames; if the similarity of the two adjacent frames is lower than the third preset threshold, determine the key frame in the two adjacent frames; determine a plot segment in the video, the plot segment includes all frames between the two adjacent plot frames, and the plot frame includes the first frame, all key frames and the last frame in the video; calculate the content similarity, the content similarity is the similarity between the plot segment and the preset information; when the content similarity is greater than the fourth preset threshold, the preset information is displayed at the second key frame.
[0275] Therefore, this solution can divide the video into plots by plot frames (the first frame of the video, the key frame, and the last frame of the video), obtain the plot segments corresponding to the same plot in the video, and display preset information similar to the plot segment at the plot frame corresponding to the plot segment, so that the displayed preset information is not abrupt relative to the plot segment, which can reduce the impact on the user's perception. The video processing method of this application does not require manpower to add preset information to the video. Therefore, this application can improve the efficiency of video processing. In order to better implement the similar content analysis of step 120 and the similar content analysis of step 220 in the video processing method, this application also provides a model for similar content analysis.
[0276] The model used for similar content analysis is a multi-task model. The model shares the network parameters of the first convolutional neural network model (CNN) to extract the basic features (deep feature map) of the input image. The basic features are connected to the feature embedding layer (embedding layer) to directly obtain the embedding1 features (embedding1 features include the first embedding vector and the second embedding vector) for multi-point positioning retrieval and recognition of the beginning and end of the film; the basic features are connected to the second convolutional network model (CNN2) and another embedding layer for further feature extraction to obtain more targeted embedding2 features (embedding2 features include the third embedding vector and the fourth embedding vector), which are used for plot segmentation. Since the recognition of the beginning and end of the film adopts the method of obtaining the same clips by temporal matching across videos, it requires the help of embedding with global image representation. Therefore, the basic features with more underlying image information are connected to the embedding1 output of the embedding layer. Since the plot segmentation requires the ability to distinguish the scenes before and after the video frame, it is necessary to abstract the scenes from the basic features. Therefore, it is necessary to further perform deep learning of CNN2 on the CNN output and obtain embedding2 with the help of another embedding layer.
[0277] (1) Model structure.
[0278] First, the CNN deep image representation. Representation 1 is embedding 1 based on the deep CNN representation. CNN 2 further performs feature selection on the CNN and then obtains embedding 2 through representation 2. The CNN deep representation module can reuse the residual neural network parameters (ResNet101 neural network parameters) pre-trained on a large-scale open source dataset (ImageNet). The structure of the ResNet101 CNN deep representation module is shown in Table 1. CNN2 reuses the fifth convolutional block (conv5) in the CNN (which is the Xth convolutional layer (conv6_x) in the sixth convolutional block, or other convolutional blocks can also be used). In this case, the initialization parameters of CNN2 can reuse the network parameters of conv5 in ResNet101. Both embedding layers here use a fully connected layer (FC) structure. You can also insert multiple FC+ReLU activation function structures in front of them. The ReLU activation function is a rectified linear unit (ReLU), also known as a rectified linear unit. It is an activation function commonly used in artificial neural networks to learn more nonlinear relationships within features and then output embeddings.
[0279] In some embodiments, the resnet101 neural network parameters can be determined according to the actual application scenario.
[0280]
[0281]
[0282] Table 1 Resnet101 feature module as CNN structure
[0283]
[0284]
[0285] Table 2 embedding1 learning branch, input is the output of Table 1
[0286]
[0287] Table 3 embedding2 branch, input is the output of Table 1
[0288] (2) Data preparation
[0289] ①. Preparation of basic triplet data:
[0290] Training requires triplet data input, so the triplet data is labeled. In the triplet consisting of an anchor (anchor, a), a positive sample (positive, p), and a negative sample (negative, n), a and p form a positive sample pair, and a and n form a negative sample pair. Multiple groups of three images can be randomly extracted from all image samples, and it is marked whether each group of three images forms a triplet, and which image a, p, and n of the triplet correspond to (for two similar images in the triplet, one can be randomly selected as a and the other as p). Note: Since the model is used to match the opening and ending segments of the same series, the opening and ending segments of the same series are usually similar, so two samples need to be extremely similar to be considered similar samples a and p. Among them, there are a total of N triplet data required for training.
[0291] ②. Scene triplet data preparation:
[0292] Preparation 1: For the above-labeled triplet data, eliminate the triplets whose positive and negative samples (or negative samples and anchor samples) belong to the same scene. For example, if the positive and negative scenes are park and park respectively - such as two different perspectives or scenic spots of the park, then eliminate this group of triplets, and finally obtain scene triplet data 1 from the basic triplets (a total of P triplets, P is less than N). At this time, we can know whether the N basic triplets are scene triplets.
[0293] Preparation 2: Extract frames or a batch of images from the application business video and annotate the scene labels of the images, such as in the woods, indoors in a traditional home, indoors in a modern home, and conference rooms. After the annotation is completed, scene triplet data 2 is generated. The generation process is as follows: randomly extract two categories (A and B) from all categories, extract a pair of images from A to form a positive sample pair, and extract an image from B to form a triplet with the positive sample pair of A. The above process is performed for each training batch for batch size (bs), where bs refers to the number of data samples captured in one training. A total of bs scene triplet data 2 is generated (a total of Q triplets. This method can generate triplets far greater than the number of triplets N and P).
[0294] (3) Training process
[0295] 1): Parameter initialization:
[0296] In the pre-training phase, conv1-conv5 use the parameters of resnet101 pre-trained on imagenet (dataset), conv6 uses the pre-trained values of conv5, and the newly added embedding layer is initialized with a Gaussian distribution with a variance of 0.01 and a mean of 0.
[0297] 2) Setting learning parameters: The learning is divided into two stages. The first stage learns all the parameters in Table 1, Table 2 and Table 3, and the second stage learns Table 3.
[0298] 3) Learning rate:
[0299] A learning rate of lr = 0.0005 is used for both training and training. After every 10 iterations, lr becomes 0.1 times of the original value.
[0300] 4) Learning process: Learning is divided into two stages, such as Figure 3a In the first stage, embedding1 is mainly trained (embedding2 is auxiliary), and the weighted sum of the two losses is calculated as the total loss (loss1). In the second stage, only embedding2 is trained (embedding1 and CNN are not updated) and only loss2 is calculated.
[0301] In the first stage, N basic triplets are iterated through epochs. An epoch is a complete dataset passed through the neural network and returned once. Each round of iteration processes all N triplets until the average epoch loss stops decreasing. (Maintaining limited learning of embedding 2 while learning embedding 1 allows the CNN to gain some awareness of the embedding 2 learning task. Limited weighted learning of embedding 2 facilitates subsequent learning of embedding 2 without affecting the learning of embedding 1.)
[0302] In some embodiments, the first stage may not learn embedding2.
[0303] In the second stage, epoch 2 iterations are performed on the Q scene triplets 2; each iteration processes the full set of Q triplets until the average epoch loss no longer decreases in a certain epoch.
[0304] 5) For each epoch, training is performed in batches. The specific operations are as follows:
[0305] (1) Take all the triplets that need to be trained in this stage (N basic triplets, or Q scene triplets2), assuming there are x triplets in total (x is N or Q), and take bs triplets as a batch, for a total of x / bs batches. Each time, take 1 batch and input it into the model update parameter (a total of x / bs times are updated to complete 1 epoch iteration).
[0306] (2) One batch of forward calculations: During training, the neural network performs forward calculations on the input triplet image to obtain embedding1 and embedding2, which are represented by e1 and e2, respectively. Both are 1x64 vectors representing floating-point features. The output is the floating-point feature representation of the triplet (e1a, e1p, e1n), (e2a, e2p, e2n).
[0307] (3) Loss calculation: Calculate loss1 and loss2. In the first stage, the weighted sum of the two is calculated to obtain the total loss. In the second stage, loss2 is used as the total loss.
[0308] (4) Model parameter update: Using stochastic gradient descent (SGD), we perform a backward gradient calculation on the loss in (3) to obtain the updated parameter values, and then update the network parameters to be learned at the corresponding stage. This completes one batch of model parameter updates.
[0309] (5) Repeat steps 2 to 4 to complete the model update for all x / bs batches.
[0310] (4) Loss
[0311] L total1 =w1L1+w2L2;
[0312] L total2 =L2;
[0313] Among them, w1 is the weight coefficient in CNN, w1=1, w2 is the weight coefficient in CNN2, w2=0.1, L1 and L2 are both triplet losses, and the formulas are as follows:
[0314] The loss function (triplet loss) is calculated for the embedding features of the triplet samples in the batch. The calculation of triplet loss is as follows, where alpha in triplet loss is the function margin, set to 0.6. Alpha is a hyperparameter used to prevent the network from outputting useless results. a is the embedding of anchor point a, X p is the embedding of the positive sample p corresponding to the anchor point a, ||X a -X p || represents the L2 distance between the embedding of anchor point a and the embedding of the positive sample p corresponding to anchor point a.
[0315] The purpose of triplet loss is to make the distance between the anchor and the negative greater than the distance from the positive by more than 0.6, where 0.6 is the value of alpha.
[0316] In some embodiments, the value of alpha is determined according to the actual application scenario.
[0317] l tri =max(||X a -X p ||-||X a -X n ||+a,0)
[0318] X n The embedding of the negative sample n corresponding to the anchor point a, ||X a -X n || represents the distance between the embedding of anchor point a and the embedding of the negative sample n corresponding to anchor point a, and a is equal to 0.6 at this time.
[0319] For the first phase:
[0320] L1: In each batch, the above formula is calculated for the embedding1 obtained by inputting the basic triples of the batch into the network, and then the average triple loss of the batch is taken as L1.
[0321] L2: In each batch, the above formula is calculated for the embedding2 obtained by inputting the network into the basic triplets of the batch that are scene triplets, and then the average triplet loss of the batch is taken as L2.
[0322] L total1 Weight the two. Since the main learning basis is embedding1, w2 is very small.
[0323] For the second phase:
[0324] L2: In each batch, the above formula is calculated by inputting the embedding2 obtained by the network into the triplets of the batch (generated from the scene triplet data 2), and then the average triplet loss of the batch is taken as the total loss.
[0325] (V) Model after training
[0326] 1) If Figure 3b As shown in the figure, the trained model is a model that can include both CNN and CNN2. CNN is used to obtain embedding1, and CNN2 is used to obtain embedding2.
[0327] 2) After training, the model is divided into two models, such as Figure 3c As shown, a model includes CNN, which is used to obtain embedding1, such as Figure 3d As shown, another model includes CNN2 to obtain embedding2.
[0328] The method described in the above embodiment will be further described below.
[0329] In this embodiment, for a certain episode that is input, the video of the episode is obtained. For example, for TV series A, there are 46 episodes, then there are 46 videos. The task of mining the opening and ending credits is to mine the opening and ending credits of each video. The method of this application mines each video separately. For each video i (i.e., the target event), 10 videos (i.e., other videos) are randomly selected from the remaining videos to form a video pair with video i, so that each video has 10 video pairs for mining. The idea of mining is to match the time periods of these 10 video pairs respectively, so that each video pair generates 0 or more matching time periods. When a time period is matched more than twice and appears at the start or end position of the video, the matching time period is the opening and ending credits of video i. This application takes the above as an example to describe the method of the embodiment of this application in detail.
[0330] like Figure 4a and Figure 4b As shown, for each video pair (i, r) above, where i represents the target video for which the opening and ending credits are to be determined, and r represents the other videos (based on the video pair composition method in the previous step, r ranges from 1 to 10), for target video i, the time segment matching algorithm needs to be performed 10 times, processing one video pair each time. The specific process of a video processing method is as follows:
[0331] (1) The preset embedding distance threshold T0=0.5, when the Euclidean distance between two embeddings is less than 0.5, it means that the two embeddings are from similar frames (ie, the preset condition of step 122 in step 120).
[0332] (2) Extract frames from the two videos in the video pair and obtain the embedding of each frame.
[0333] In some embodiments, there are multiple ways to extract frames, for example, one frame may be extracted every 1 second of the video, one frame may be extracted every 2 seconds of the video, one frame may be extracted every 10 seconds of the video, and so on.
[0334] (III) Frame-level similarity matching (frame matching). For each frame j (i.e., target frame) in video i: calculate its Euclidean distance with the embedding of each frame in video r, and take frame j as a similar frame to other frames less than T0, obtain the list of other frames (or matching frames) corresponding to j as similar frames (sim-id-list), and record the corresponding similar frame time deviation diff-time-list (e.g., for frame j = 1, sim-id-list is [1, 2, 3], indicating that it is similar to the 1st, 2nd, and 3rd seconds of video r; the frame sequence difference diff-time-list is [0, 1, 2], indicating the distance between the time represented by other frames in sim-id-list and frame j = 1 (i.e., frame sequence difference). Here, the default frame extraction is 1 frame per second, so the frame number is the number of seconds).
[0335] In some embodiments, if the frame extraction is to extract a frame every predetermined time period in the video, the time offset is equal to the frame sequence difference multiplied by the predetermined time period.
[0336] (4) Traverse all frames and count the number of matching frames between video i and video r (i.e., the number of matching frames between j and r in step 3). If the number of matching frames is less than 1, then video i and r do not have the same video segment, and no opening or ending credits can be mined. Otherwise, proceed to the next step.
[0337] (5) dt is reordered to obtain the SL list: all matching frames in SL are sorted from small to large according to diff-time (i.e., dt). When dt is the same, they are sorted from small to large according to the frame number of the target frame of video i in SL, and the corresponding diff-time-list is reorganized in this order.
[0338] For example, the frames with a difference of 0 are placed at the front, and those with a difference of 1 are placed at the back, and so on. For example, the new SL list is [10,11], [11,12], [2,4], [3,5], [4,6], [6,9], [7,10]. The numbers before "," refer to the target frames in video i, and the numbers after "," refer to other frames in video r. The target frames before "," are similar frames to the other frames after ",".
[0339] (6) Frame matching is merged into segment matching based on the same frame sequence difference.
[0340] Reorganize the data using dt to obtain match-dt-list: For the lists in the similar frame list SL of all frames of video i, reorganize them using the frame sequence difference as the primary key to obtain a list of dt from small to large, and obtain the similar frames match-dt-list with frame sequence difference of 0, 1, 2, etc.: {0:{count,start-id,match-id-list},…}, for example {2:{3,2,[[2,4],[3,5],[4,6]]}, 3:{2,6,[[6,9],[7,10]]}}, where 2 refers to the time difference of 2. For example, the second frame of video i and the fourth frame of video r are similar, then the time difference between the two frames is 2; count is the number of similar frames under this time deviation. For example, if the second frame of video i and the fourth frame of video r are similar, then count is increased by 1; start-id refers to the minimum frame ID of similar frames under the same frame sequence difference. For example, if the target frame with frame sequence number 2 of video i is similar to another frame with frame sequence number 4, then start-id is 2. (VII) First merge segment. Merge the two dt lists in the match-dt-list whose preceding and following dts are less than 3 (i.e., merge the matching pairs whose frame sequence difference is within 3 seconds). Merge the one with the larger dt into the one with the smaller dt. Simultaneously, update the matching frames with the larger dts. Simultaneously, update the matching frame list SL in step 5.
[0341] For example, as in the above example, the frame sequence difference diff-time-list is [1,2,3], then the larger dt in the list is 3, and the smaller dt is 2, that is, [2,4], [3,5], [4,6] with dt of 2 (i.e., the first similar segment) and [6,9], [7,10] with dt of 3 (i.e., the second similar segment) can be merged, and finally {2:{5, 2, [[2,4], [3,5], [4,6], [6,8], [7,9]]}} (i.e., the first merged segment) is obtained, where count is the sum of counts of dt = 2 and dt = 3. start-id finds the frame with the smallest i video from the similar frame lists of dt = 2 and dt = 3. For the list of dt = 3, the frame numbers of other frames corresponding to the similar frames are rewritten, such as rewriting [6,9] to [6,8] and merging it into the similar frame list of dt = 2. At the same time, the similar frame pairs with rewritten frame numbers are synchronously updated to the SL matching frame list in step five, such as: [10,11], [11,12], [2,4], [3,5], [4,6], [6,8], [7,9].
[0342] (8) Since the merged frame list mentioned above may disrupt the order of dt or frame id, it is necessary to re-sort dt, that is, perform the sorting of step 5 again on the new SL list to obtain the sorted matching frame list.
[0343] (9) Reorganize the data using dt to obtain match-dt-list: Execute step 6 again.
[0344] (10) Calculate the match-duration-list:
[0345] A1. The time interval between the two matching segments is preset to be greater than T2.
[0346] For example, if T2 is 8 seconds and there is 1 frame per second, the frame numbers differ by 8.
[0347] A2. For each dt in match-dt-list (e.g., dt=2):
[0348] B1. For each frame srcT of video i under dt (such as 2 in the above examples 2, 3, 4, 6, and 7):
[0349] If the difference between C1 and srcT is greater than T2 (e.g., the difference between 2 and the previous srcT is 9, which is greater than the interval threshold), the previous similar frame pairs are merged into one matching segment, and new similar frame pair statistics are started from the current srcT, and the similar frames are stored in a temporary list tmplist. If dt = 2 and srcT = 2, similar frames in the previous temporary frame list are saved as matching segments. For example, similar frames from the previous tmplist = [[10,11],[11,12]] are added to the match-duration-list as matching segments. For example, matching segment information like [10,11,11,12,1,2,2] is added, where each value represents [src-startTime, src-endTime, ref-startTime, ref-endTime, dt, duration, count]. This means that the matching segment stores two video segments: the start and end frames of video i, the start and end frames of the matching video, the dt of the matching segment, the duration of the matching segment, and the number of similar frames matched. The similar frames this time are saved in the temporary list tmplist = [[2,4]].
[0350] C2. When the difference between srcT and the previous srcT is less than T2, the similar frames are stored in a temporary list tmplist. For example, for dt2, srcT = 3, 4, 6, and 7 are all stored in the temporary list, resulting in tmplist = [[2, 4], [3, 5], [4, 6], [6, 8], [7, 9]]. When the current dt is the last similar frame (for example, srcT = 7), the accumulated similar frames in tmplist form a matching segment and are added to match-duration-list. For example, [2, 7, 4, 9, 2, 6, 5] is added, where the duration is 7-2+1 and count = 5 is the similar frame count. Thus, match-duration-list = [[10, 11, 11, 12, 1, 2, 2], [2, 7, 4, 9, 2, 6, 5]].
[0351] (11) Sort the above match-duration-list in reverse order by the number of similar frames count, such as match-duration-list = [[2,7,4,9,2,6,5], [10,11,11,12,1,2,2]].
[0352] (12) Second merged segment. Process the overlapping time periods in the match-duration-list. Since similarity calculation involves traversing all frames of the two videos and performing distance calculations to find similarities within a certain threshold range, it is easy for a frame to be similar to multiple frames, resulting in two matching time periods in the match-duration-list overlapping. This situation needs to be handled.
[0353] A1. Set the minimum matching duration T3 (e.g. 5, indicating a minimum matching duration of 5 seconds).
[0354] A2. For time period i in match-duration-list (the time period consisting of src-startTime and src-endTime):
[0355] B1. For time period j=i+1 in match-duration-list, time period j (i.e., the position of the fourth similar segment in the target video i) refers to the time period in match-duration-list that is adjacent to time period i (i.e., the position of the third similar segment in the target video i).
[0356] C1, such as Figure 4c As shown in 1, when time period i contains time period j, time period j is deleted.
[0357] C2, such as Figure 4cAs shown in 2, when time period i and time period j have an intersection and the starting point of time period i is the earliest starting point, the starting point of time period j is moved back to the end position of i, and time period j is updated (that is, the position of the second merged segment in the target video i). At this time, when the length of time period j is less than T3, time period j is deleted, otherwise the new time period j replaces the old time period j.
[0358] C3, such as Figure 4c As shown in 3, when time periods i and j intersect and the starting point of time period j is the earliest, the end point of time period j is moved forward to the starting point of time period i and time period j is updated. At this time, if the length of the updated time period j is less than T3, j is deleted. Otherwise, the new time period j replaces the old one.
[0359] (13) Return matching time period information, such as match-duration-list = [[2,7,4,9,2,6,5], [10,11,11,12,1,2,2]], or only return matching segments [[2,7,4,9], [10,11,11,12]]; otherwise, matching segments include similar segments composed of similar frames. For example, the similar segments in matching segment [[2,7,4,9] are the segments corresponding to frame numbers 2 to 7 in target video i.
[0360] (14) For video i, it is mined from other videos vid2, other videos vid3, and other videos vid4. Then, for a total of N = 3 video pairs [I, vid2][I, vid3], [I, vid4], the video segment matching from step 1 to step 13 is performed respectively, and 3 matching information is obtained. For example, the matching segments of the first pair of videos are returned: [[2, 7, 4, 9], [10, 11, 11, 12]], the matching segments of the second pair of videos are returned [[2, 7, 4, 9]], and the matching segments of the third pair of videos are returned [[2, 7, 4, 10]].
[0361] (15) Count the number of matching segments, such as [2,7,4,9] has 2 times, [2,7,4,10] has 1 time, and [10,11,11,12] has 1 time.
[0362] (16) Sort the matching segments in reverse order by count. If the counts are the same, sort them in ascending order by src-startTime: match-list = [[2,7,4,9], [2,7,4,10], [10,11,11,12]], count-list = [2,1,1].
[0363] (17) Merge overlapping matching segments in the match-list:
[0364] A1. Set the effective overlap ratio T4 (e.g., 0.5, indicating that when the intersection duration of two time segments is greater than the target segment duration, the counts of the two segments need to be combined), and the effective match count T5 (e.g., 3, indicating that when the similarity frame count in a matching segment is greater than T5, the segment cannot be ignored).
[0365] A2. For time period i in the match-list, where time period i is the position of the third similar segment in target video i) (referring to the time period composed of src-startTime and src-endTime):
[0366] B1. For time period j=i+1 in the match-list, time period j is the position of the fourth similar segment in the target video i. Time period j is the time period adjacent to time period i in the match-list:
[0367] C1, such as Figure 4c As shown in 1, when time period i contains time period j and the length of time period j is greater than 0.5*the length of time period i, time period j is deleted.
[0368] C2, when time i and time j have an intersection, when the intersection duration is greater than 0.5*i segment duration, the intersection duration is the position of the overlapping segment between the third similar segment and the fourth similar segment in the target video i:
[0369] D1, such as Figure 4c As shown in 2 and 3, when the number of similar frames in segment j is greater than T5, the merged time segment i and time segment j is the longest start and end time.
[0370] D2: When the number of similar frames in segment j is less than T5, time j is deleted. (That is, segments i and j are not merged at this time, only the segment i with the most occurrences is retained, but the number of occurrences of segment j is reflected in the new segment i count).
[0371] C3. When i and j intersect, and the duration of the intersection is less than 0.5*the duration of segment i, discard segment j.
[0372] (XVIII) Obtain a new video matching segment match-list (such as [[2,7,4,9], [10,11,11,12]]) and a count count-list (such as [3,1]). The count of the matching segment in the count-list is equal to the number of similar segments.
[0373] (19) Setting a valid recurrence ratio threshold T6 indicates that in the mining of N video pairs, when the recurrence times x of a matching video segment is greater than N*T6, it is a valid repeated segment (eg, T6=0.5).
[0374] (20) Based on the quantity, determine the target segment among all similar segments.
[0375] For example, the number of matching segments [2, 7, 4, 9] is the largest, and the segments corresponding to frame numbers 2 to 7 in [2, 7, 4, 9] in the target video i are taken as target segments.
[0376] (21) Based on the target segment, determine the key frames in the target video.
[0377] In some embodiments, determining a key frame in a target video based on a target segment includes:
[0378] Determining at least one transition frame from the target video, where the transition frame includes text and a preset background;
[0379] determining a target transition frame from at least one transition frame, wherein the target transition frame is adjacent to the target segment;
[0380] Merging all intermediate frames, target transition frames, and target segments into a new target segment, where the intermediate frames are frames between the target transition frames and the target segment;
[0381] Based on the new target segment, key frames are identified in the target video.
[0382] For example, a video frame is identified as black text based on a classification model (a black text binary classification model needs to be pre-trained to identify whether an image is black text, as shown below). For example, in addition to the match-list [[2,7,4,9]] above, [11,12,14,15] and [30,31,32] are identified as other images of black text. From all black text segments, find the one closest to the previously retrieved title end time. If the start time of the black text segment is less than T7 from the title end time (such as 5, indicating that the black text appears within 5 seconds of the title end), it indicates that the text is the announcement before the main film begins, and it is merged into the target segment.
[0383] In some embodiments, determining a target segment among all the similar segments based on the number includes:
[0384] Obtaining a preset position of a preset segment in a video, where each video in the video collection includes at least a portion of the preset segment;
[0385] Based on the quantity, candidate segments are determined among all similar segments;
[0386] Comparing the position of the target segment in the target video with the preset position to obtain the distance between the candidate segment and the preset segment;
[0387] According to the distance, a target segment is determined from multiple candidate segments.
[0388] For example, Figure 4d As shown in the figure, since the producer's title is generally fixed, this paper uses embedding1 as the feature time period matching to locate the title position. The specific process is to first collect the producer's clips and put them into the inventory, obtain the preset clips from the inventory, and then form a video pair for each candidate clip and the video in the inventory, perform time period matching, and find the candidate clip closest to the preset clip as the target clip.
[0389] In some embodiments, determining a key frame in a target video based on the target segment includes:
[0390] Obtaining preset text in a target video, where the preset text is associated with a target frame in the target video;
[0391] Determining target text from preset texts, where the target text is used to indicate the playback order of the target video in the video collection;
[0392] According to the target text, a key frame is determined in the target video, where the key frame is a target frame associated with the target text.
[0393] In some embodiments, based on the video's dialogue file, the location of the words "episode n" (i.e., episode identification) can be found. This can then be used to locate the episode's location (i.e., key frame) in the main film. Furthermore, black screen text can be identified to determine the frame where the episode begins. This can then provide a time location closer to the main film where an ad can be inserted.
[0394] From the above, it can be seen that by using the technology of video frame retrieval and the frame sequence matching between multiple videos for positioning the beginning and ending of the video in this solution, the identification and positioning of the beginning and ending of the video can be achieved when the time is not aligned or the beginning and ending of the video are of unequal length. Since the target segment appears repeatedly, it can be seen that the target segment will not affect the content other than the target segment in the target video. The key frame in the target video can be determined through the target segment. The key frame is the turning point of the content in the target video. By displaying preset information at the position of the key frame, the impact on the user's perception can be reduced. The video processing method of the present application can quickly determine the turning point of the content in the target video, and there is no need to consume manpower to determine the turning point of the content during the video viewing process. Therefore, the present application improves the efficiency of video processing.
[0395] In this embodiment, if Figure 4e As shown, the plot segmentation of the video is taken as an example. This application takes the above as an example to explain the method of the embodiment of the present application in detail. The specific process of a video processing method is as follows:
[0396] (1) Extract frames from the video according to preset rules to obtain frame-level images, obtain the embedding2 features of each frame, and based on whether the Euclidean distance between the embedding2 of the previous and next frames is less than the preset threshold 1 (thr1), cluster the previous and next frames similarly to preliminarily determine whether the previous and next frames are from the same plot;
[0397] (2) Merge the plot segments. Starting from the first plot segment, if there is a sufficiently similar plot segment, merge the two plot segments (and the plot in between) to obtain the secondary plot segmentation of the video. For the original video plot, the positions of all plot segments are obtained from the second plot segment to the last plot segment according to the following process: determine the similarity between its plot embedding2 and the previous plot embedding2 (calculate the Euclidean distance between each two frames of the two plots, where the number of frames with a distance less than the preset threshold 1 is divided by the smallest number of frames in the two plot frames); if it is greater than the preset threshold 2, it is determined to be the same plot; if it is less than, start a new plot.
[0398] (3) Obtain the video script, segment the previous plot, and if the segmentation time point is between lines, move the time point back to include the entire line.
[0399] For example, when there are lines within 2 seconds before and after a certain split time point, the split point moves to after the next line.
[0400] (4) To improve the matching effect, the embedding2 of each plot can be recorded, and then the embedding2 of each advertising video frame can be obtained; according to the above-mentioned plot merging method, for each plot segment, the advertisement with the highest similarity among all advertisements is found, and the advertisement is inserted after the plot segment.
[0401] As can be seen from the above, by using plot measurement features to compare and aggregate the previous and next frames of the video, plot segmentation is achieved to obtain the key frames in the target video. In this way, displaying preset information at the position of the key frames can reduce the impact on the user's viewing experience. The video processing method of this application can quickly determine the content turning point in the target video, and there is no need to spend manpower to determine the content turning point during the video viewing process. Therefore, this application improves the efficiency of video processing.
[0402] In order to better implement the above method, an embodiment of the present application further provides a video processing device, which can be specifically integrated into an electronic device, and the electronic device can be a terminal, a server, or other device.
[0403] For example, Figure 5As shown, the video processing apparatus may include a first acquisition unit 510, a first analysis unit 520, a quantity determination unit 530, a segment determination unit 540, and a first target determination unit 550, as follows:
[0404] (1) First acquisition unit 510.
[0405] The first acquisition unit 510 is configured to acquire a video set, where the multiple videos in the video set include a target video and at least one other video.
[0406] (2) First analysis unit 520.
[0407] The first analyzing unit 520 is configured to perform similar content analysis on the target video and other videos to obtain similar segments between the target video and other videos, as well as positions of the similar segments in the target video.
[0408] In some embodiments, a target video includes a target frame set, the target frame set includes multiple target frames and a frame sequence number of each target frame, and other videos include other frame sets, the other frame sets include multiple other frames. Similar content analysis is performed on the target video and other videos to obtain similar segments in the target video and other videos, as well as positions of the similar segments in the target video, including:
[0409] Calculate the similarity between the target frame and other frames;
[0410] When the similarity meets the preset conditions, the target frame is used as the similar frame of other frames, and the frame number of the target frame is used as the frame number of the similar frame;
[0411] Determine at least one similar segment from all similar frames, where the similar segment includes at least two similar frames, and the frame sequence numbers of the at least two similar frames are continuous;
[0412] According to the frame sequence number of each similar frame in the similar segment, the position of the similar segment in the target video is determined.
[0413] In some embodiments, the other frame sets further include a frame sequence number of each other frame, and determining at least one similar segment from all similar frames includes:
[0414] Determine a frame sequence difference, where the frame sequence difference is the difference between the frame sequence number of the similar frame and the frame sequence number of the corresponding other frames;
[0415] At least one similar segment is determined from all similar frames corresponding to the same frame sequence difference.
[0416] In some embodiments, the similar segments include a first merged segment, and after determining at least one similar segment from all similar frames corresponding to the same frame sequence difference value, the method further includes:
[0417] Determine a first difference, where the first difference is an absolute value of a difference between the first frame sequence difference and the second frame sequence difference, the first frame sequence difference is any one of a plurality of frame sequence differences, and the second frame sequence difference is a frame sequence difference other than the first frame sequence difference;
[0418] When the first difference is not greater than a first preset threshold, determining a second difference, where the second difference is an absolute value of a difference between a frame sequence number of a first frame in the first similar segment and a frame sequence number of a second frame in the second similar segment, the first similar segment is a similar segment corresponding to the first frame sequence difference, the second similar segment is a similar segment corresponding to the second frame sequence difference, and the first frame is adjacent to the second frame;
[0419] When the second difference is not greater than a second preset threshold, the first similar segment and the second similar segment are merged to obtain a first merged segment.
[0420] (3) Quantity determination unit 530.
[0421] The number determining unit 530 is configured to determine the number of similar segments at each position.
[0422] In some embodiments, the similar segments include the second merged segment, and determining the number of similar segments at each position includes:
[0423] determining, according to positions of the third similar segment and the fourth similar segment in the target video, an overlapping segment between the third similar segment and the fourth similar segment, where the third similar segment is any one of a plurality of similar segments, the fourth similar segment is a similar segment other than the third similar segment, and the plurality of similar segments include similar segments in the target video and each of the other videos;
[0424] Merging the third similar segment and the fourth similar segment according to the overlapping segments to obtain a second merged segment;
[0425] The number of similar segments at each position is determined, where the similar segments include the second merged segment and unmerged segments, and the unmerged segments are similar segments excluding the third similar segment and the fourth similar segment.
[0426] (4) Segment determination unit 540.
[0427] The segment determination unit 540 is configured to determine a target segment among all similar segments based on the quantity.
[0428] In some embodiments, determining a target segment among all the similar segments based on the number includes:
[0429] Obtaining a preset position of a preset segment in a video, where each video in the video collection includes at least a portion of the preset segment;
[0430] Based on the quantity, candidate segments are determined among all similar segments;
[0431] Comparing the position of the target segment in the target video with the preset position to obtain the distance between the candidate segment and the preset segment;
[0432] According to the distance, a target segment is determined from multiple candidate segments.
[0433] (5) First target determination unit 550.
[0434] The first target determination unit 550 is configured to determine a key frame in a target video based on the target segment, so as to display preset information at a position of the key frame in the target video.
[0435] In some embodiments, determining a key frame in a target video based on a target segment includes:
[0436] Determining at least one transition frame from the target video, where the transition frame includes text and a preset background;
[0437] determining a target transition frame from at least one transition frame, wherein the target transition frame is adjacent to the target segment;
[0438] Merge all intermediate frames, target transition frames, and target segments to obtain new target segments, where the intermediate frames are the frames between the target transition frames and the target segments;
[0439] Based on the new target segment, key frames are identified in the target video.
[0440] In some embodiments, determining a key frame in a target video based on a target segment includes:
[0441] Obtaining preset text in a target video, where the preset text is associated with a target frame in the target video;
[0442] Determining target text from preset texts, where the target text is used to indicate the playback order of the target video in the video collection;
[0443] According to the target text, a key frame is determined in the target video, where the key frame is a target frame associated with the target text.
[0444] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.
[0445] As can be seen from the above, the video processing device of this embodiment obtains a video set by the first acquisition unit, and the multiple videos in the video set include a target video and at least one other video; the first analysis unit performs similarity content analysis on the target video and the other videos to obtain similar segments in the target video and the other videos, as well as the positions of the similar segments in the target video; the quantity determination unit determines the number of similar segments at each position; the segment determination unit determines the target segment among all similar segments based on the quantity; the first target determination unit determines the key frame in the target video based on the target segment, so as to display preset information at the position of the key frame of the target video.
[0446] Therefore, the video processing method of the present application can quickly determine the content turning points in the target video, and there is no need to spend manpower to determine the content turning points during the video viewing process. Therefore, the present application improves the efficiency of video processing.
[0447] For example, Figure 6 As shown, the video processing device may further include a second acquisition unit 610, a second analysis unit 620, a second target determination unit 630, a plot determination unit 640, a similarity calculation unit 650, and a presentation unit 660, as follows:
[0448] (1) Second acquisition unit 610.
[0449] The second acquiring unit 610 is configured to acquire a video.
[0450] (2) The second analysis unit 620.
[0451] The second analysis unit 620 is configured to perform similarity analysis on two adjacent frames in the video to obtain a similarity between the two adjacent frames.
[0452] (3) Second target determination unit 630.
[0453] The second target determination unit 630 is configured to determine a key frame between the two adjacent frames if the similarity between the two adjacent frames is lower than a third preset threshold.
[0454] In some embodiments, if the similarity between two adjacent frames is lower than a third preset threshold, determining a key frame between the two adjacent frames includes:
[0455] Get the preset sentence corresponding to the video;
[0456] If the similarity between the two adjacent frames is lower than a third preset threshold, performing content recognition processing on the audio content corresponding to each video frame in the two adjacent frames to obtain recognition text corresponding to the two adjacent frames;
[0457] Determine, from the preset sentences, a target sentence that is identical to the recognition text corresponding to the two adjacent frames;
[0458] According to the target sentence, the key frame is determined between two adjacent frames.
[0459] In some embodiments, determining a key frame between two adjacent frames according to a target sentence includes:
[0460] When the target sentence is adjacent to the preset symbol, one of the two adjacent frames corresponding to the target sentence is used as the key frame.
[0461] In some embodiments, determining a key frame between two adjacent frames according to a target sentence includes:
[0462] When the target sentence is not adjacent to the preset symbol, content recognition processing is performed on the audio content corresponding to other video frames in the video to obtain recognized text corresponding to the other video frames, where the other video frames are the video frames after the two adjacent frames in the video;
[0463] Determining other sentences in the preset sentences that are identical to the recognized texts corresponding to the other video frames;
[0464] When other sentences are adjacent to the preset symbol, other video frames corresponding to the other sentences are used as key frames.
[0465] (4) Plot determination unit 640.
[0466] The plot determination unit 640 is configured to determine a plot segment in the video. The plot segment includes all frames between two adjacent plot frames. The plot frames include the first frame, all key frames, and the last frame in the video.
[0467] (5) Similarity calculation unit 650.
[0468] The similarity calculation unit 650 is used to calculate content similarity, where the content similarity is the similarity between the plot segment and the preset information.
[0469] (6) Display unit 660.
[0470] The display unit 660 is configured to display the preset information at the plot frame corresponding to the plot segment when the content similarity is greater than a fourth preset threshold.
[0471] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.
[0472] As can be seen from the above, the video processing device of this embodiment obtains the video by the second acquisition unit; the second analysis unit performs similar content analysis on two adjacent frames in the video to obtain the similarity between the two adjacent frames; if the similarity between the two adjacent frames is lower than the third preset threshold, the second target determination unit determines the key frame in the two adjacent frames; the plot determination unit determines the plot segment in the video, the plot segment includes all frames between two adjacent plot frames in the plot frame set, and the plot frames in the plot frame set include the first frame, all key frames and the last frame in the video; the similarity calculation unit calculates the content similarity, which is the similarity between the plot segment and the preset information; when the content similarity is greater than the fourth preset threshold, the display unit displays the preset information at the plot frame corresponding to the plot segment.
[0473] Therefore, the video processing method of the present application can quickly determine the content turning points in the target video, and there is no need to spend manpower to determine the content turning points during the video viewing process. Therefore, the present application improves the efficiency of video processing.
[0474] An embodiment of the present application also provides an electronic device, which may be a terminal, a server, or other device.
[0475] In this embodiment, the electronic device of this embodiment is a server as an example for detailed description, for example, Figure 7 As shown, it shows a schematic diagram of the structure of the server involved in the embodiment of the present application, specifically:
[0476] The server may include one or more processing core processors 710, one or more computer-readable storage media memories 720, a power supply 730, an input module 740, and a communication module 750. It will be understood by those skilled in the art that Figure 7 The server structure shown in the figure does not constitute a limitation on the server, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0477] Processor 710 is the server's control center, connecting various components of the server using various interfaces and circuits. It executes software programs and / or modules stored in memory 720 and accesses data stored in memory 720 to perform various server functions and process data. In some embodiments, processor 710 may include one or more processing cores. In some embodiments, processor 710 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 710.
[0478] The memory 720 can be used to store software programs and modules. The processor 710 executes various functional applications and data processing by running the software programs and modules stored in the memory 720. The memory 720 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created based on the use of the server, etc. In addition, the memory 720 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 720 may also include a memory controller to provide the processor 710 with access to the memory 720.
[0479] The server also includes a power supply 730 that supplies power to various components. In some embodiments, the power supply 730 can be logically connected to the processor 710 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 730 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0480] The server may further include an input module 740, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0481] The server may also include a communication module 750. In some embodiments, the communication module 750 may include a wireless module. The server may use the wireless module of the communication module 750 to perform short-range wireless transmission, thereby providing users with wireless broadband Internet access. For example, the communication module 750 may be used to help users send and receive emails, browse web pages, and access streaming media.
[0482] From the above, it can be seen that the two video processing methods of this application can quickly determine the content turning points in the target video, and there is no need to spend manpower to determine the content turning points during the video viewing process. Therefore, this application improves the efficiency of video processing.
[0483] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0484] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the video processing methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:
[0485] A video processing method, comprising:
[0486] Obtain a video collection, where the multiple videos in the video collection include a target video and at least one other video;
[0487] Analyze the target video and other videos for similar content, and obtain similar segments in the target video and other videos, as well as the positions of the similar segments in the target video;
[0488] Determine the number of similar fragments at each position;
[0489] Based on the quantity, the target fragment is determined among all similar fragments;
[0490] Based on the target segment, a key frame is determined in the target video so as to display preset information at a position of the key frame of the target video.
[0491] Another video processing method includes:
[0492] Get video and preset information;
[0493] Perform similar content analysis on two adjacent frames in the video to obtain the similarity between the two adjacent frames;
[0494] If the similarity between the two adjacent frames is lower than a third preset threshold, determining a key frame between the two adjacent frames;
[0495] Determine a plot segment in the video, where the plot segment includes all frames between two adjacent plot frames, and the plot frames include the first frame, all key frames, and the last frame in the video;
[0496] Calculate content similarity, which is the similarity between the plot segment and the preset information;
[0497] When the content similarity is greater than a fourth preset threshold, the preset information is displayed at the plot frame corresponding to the plot segment.
[0498] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0499] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in the various optional implementations of the video processing aspects provided in the above-described embodiments.
[0500] Since the instructions stored in the storage medium can execute the steps in any video processing method provided in the embodiments of the present application, the beneficial effects that can be achieved by any video processing method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0501] The above is a detailed introduction to a video processing method, device, server and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A video processing method, characterized in that: include: Acquire a video set, where the multiple videos in the video set include a target video and at least one other video; Performing similar content analysis on the target video and the other videos to obtain similar segments in the target video and the other videos, as well as positions of the similar segments in the target video; determining the number of said similar segments at each of said locations; determining a target segment among all the similar segments based on the number; Based on the target segment, a key frame is determined in the target video, so as to display preset information at a position of the key frame in the target video.
2. The video processing method according to claim 1, wherein: The target video includes a target frame set, the target frame set includes multiple target frames and a frame sequence number of each target frame, the other video includes another frame set, the other frame set includes multiple other frames, and performing similar content analysis on the target video and the other video to obtain similar segments in the target video and the other video, as well as positions of the similar segments in the target video, includes: Calculating the similarity between the target frame and the other frames; When the similarity meets a preset condition, the target frame is used as a similar frame of the other frames, and the frame sequence number of the target frame is used as the frame sequence number of the similar frame; Determine at least one similar segment from all the similar frames, wherein the similar segment includes at least two similar frames, and the frame sequence numbers of the at least two similar frames are continuous; The position of the similar segment in the target video is determined according to the frame sequence number of each similar frame in the similar segment.
3. The video processing method according to claim 2, wherein: The other frame set further includes a frame sequence number of each of the other frames, and the determining at least one similar segment from all the similar frames includes: Determine a frame sequence difference value, where the frame sequence difference value is a difference between a frame sequence number of the similar frame and a frame sequence number of the corresponding other frame; At least one similar segment is determined from all similar frames corresponding to the same frame sequence difference.
4. The video processing method according to claim 3, wherein: The similar segments include a first merged segment, and after determining at least one similar segment from all similar frames corresponding to the same frame sequence difference value, the method further includes: Determine a first difference, where the first difference is an absolute value of a difference between a first frame sequence difference and a second frame sequence difference, the first frame sequence difference is any one of a plurality of the frame sequence differences, and the second frame sequence difference is a frame sequence difference other than the first frame sequence difference; When the first difference is not greater than a first preset threshold, determining a second difference, where the second difference is an absolute value of a difference between a frame sequence number of a first frame in a first similar segment and a frame sequence number of a second frame in a second similar segment, the first similar segment being the similar segment corresponding to the first frame sequence difference, the second similar segment being the similar segment corresponding to the second frame sequence difference, and the first frame being adjacent to the second frame; When the second difference is not greater than a second preset threshold, the first similar segment and the second similar segment are merged to obtain a first merged segment.
5. The video processing method according to claim 1, wherein: The similar segments include a second merged segment, and determining the number of the similar segments at each position includes: determining, based on the positions of the third similar segment and the fourth similar segment in the target video, an overlapping segment between the third similar segment and the fourth similar segment, where the third similar segment is any one of a plurality of similar segments, and the fourth similar segment is a similar segment other than the third similar segment, where the plurality of similar segments include similar segments in the target video and each of the other videos; Merging the third similar segment and the fourth similar segment according to the overlapping segments to obtain a second merged segment; The number of the similar segments at each of the positions is determined, where the similar segments include the second merged segment and unmerged segments, and the unmerged segments are similar segments other than the third similar segment and the fourth similar segment.
6. The video processing method according to claim 1, wherein: The determining of key frames in the target video based on the target segment includes: Determining at least one transition frame from the target video, wherein the transition frame includes text and a preset background; determining a target transition frame from the at least one transition frame, wherein the target transition frame is adjacent to the target segment; Merging all intermediate frames, the target transition frame, and the target segment to obtain a new target segment, wherein the intermediate frame is a frame between the target transition frame and the target segment; Based on the new target segment, key frames are determined in the target video.
7. The video processing method according to claim 1, wherein: Determining a target segment among all the similar segments based on the quantity includes: Obtaining a preset position of a preset segment in the video, each of the videos in the video set including at least part of the preset segment; determining a candidate segment among all the similar segments based on the number; Comparing the position of the candidate segment in the target video with the preset position to obtain the distance between the candidate segment and the preset segment; A target segment is determined from the plurality of candidate segments according to the distance.
8. The video processing method according to claim 1, wherein: The determining of key frames in the target video based on the target segment includes: Acquire preset text in the target video, where the preset text is associated with a target frame in the target video; Determining target text from the preset text, where the target text is used to indicate the playback order of the target video in the video collection; According to the target text, a key frame is determined in the target video, where the key frame is the target frame associated with the target text.
9. A video processing device, characterized in that: include: A first acquisition unit is configured to acquire a video set, where the multiple videos in the video set include a target video and at least one other video; A first analyzing unit is configured to perform similar content analysis on the target video and the other videos to obtain similar segments in the target video and the other videos, as well as positions of the similar segments in the target video; a quantity determining unit, configured to determine the quantity of the similar segments at each of the locations; a segment determining unit, configured to determine a target segment among all the similar segments based on the number; The first target determining unit is configured to determine a key frame in the target video based on the target segment, so as to display preset information at a position of the key frame in the target video.
10. The video processing device according to claim 9, wherein: The target video includes a target frame set, the target frame set includes multiple target frames and a frame sequence number of each target frame, the other video includes another frame set, the other frame set includes multiple other frames, and performing similar content analysis on the target video and the other video to obtain similar segments in the target video and the other video, as well as positions of the similar segments in the target video, includes: Calculating the similarity between the target frame and the other frames; When the similarity meets a preset condition, the target frame is used as a similar frame of the other frames, and the frame sequence number of the target frame is used as the frame sequence number of the similar frame; Determine at least one similar segment from all the similar frames, wherein the similar segment includes at least two similar frames, and the frame sequence numbers of the at least two similar frames are continuous; The position of the similar segment in the target video is determined according to the frame sequence number of each similar frame in the similar segment.
11. The video processing device according to claim 10, wherein: The other frame set further includes a frame sequence number of each of the other frames, and the determining at least one similar segment from all the similar frames includes: Determine a frame sequence difference value, where the frame sequence difference value is a difference between a frame sequence number of the similar frame and a frame sequence number of the corresponding other frame; At least one similar segment is determined from all similar frames corresponding to the same frame sequence difference.
12. The video processing device according to claim 11, wherein: The similar segments include a first merged segment, and after determining at least one similar segment from all similar frames corresponding to the same frame sequence difference value, the method further includes: Determine a first difference, where the first difference is an absolute value of a difference between a first frame sequence difference and a second frame sequence difference, the first frame sequence difference is any one of a plurality of the frame sequence differences, and the second frame sequence difference is a frame sequence difference other than the first frame sequence difference; When the first difference is not greater than a first preset threshold, determining a second difference, where the second difference is an absolute value of a difference between a frame sequence number of a first frame in a first similar segment and a frame sequence number of a second frame in a second similar segment, the first similar segment being the similar segment corresponding to the first frame sequence difference, the second similar segment being the similar segment corresponding to the second frame sequence difference, and the first frame being adjacent to the second frame; When the second difference is not greater than a second preset threshold, the first similar segment and the second similar segment are merged to obtain a first merged segment.
13. The video processing device according to claim 9, wherein: The similar segments include a second merged segment, and determining the number of the similar segments at each position includes: determining, based on the positions of the third similar segment and the fourth similar segment in the target video, an overlapping segment between the third similar segment and the fourth similar segment, where the third similar segment is any one of a plurality of similar segments, and the fourth similar segment is a similar segment other than the third similar segment, where the plurality of similar segments include similar segments in the target video and each of the other videos; Merging the third similar segment and the fourth similar segment according to the overlapping segments to obtain a second merged segment; The number of the similar segments at each of the positions is determined, where the similar segments include the second merged segment and unmerged segments, and the unmerged segments are similar segments other than the third similar segment and the fourth similar segment.
14. The video processing device according to claim 9, wherein: The determining of key frames in the target video based on the target segment includes: Determining at least one transition frame from the target video, wherein the transition frame includes text and a preset background; determining a target transition frame from the at least one transition frame, wherein the target transition frame is adjacent to the target segment; Merging all intermediate frames, the target transition frame, and the target segment to obtain a new target segment, wherein the intermediate frame is a frame between the target transition frame and the target segment; Based on the new target segment, key frames are determined in the target video.
15. The video processing device according to claim 9, wherein: Determining a target segment among all the similar segments based on the quantity includes: Obtaining a preset position of a preset segment in the video, each of the videos in the video set including at least part of the preset segment; determining a candidate segment among all the similar segments based on the number; Comparing the position of the candidate segment in the target video with the preset position to obtain the distance between the candidate segment and the preset segment; A target segment is determined from the plurality of candidate segments according to the distance.
16. The video processing device according to claim 9, wherein: The determining of key frames in the target video based on the target segment includes: Acquire preset text in the target video, where the preset text is associated with a target frame in the target video; Determining target text from the preset text, where the target text is used to indicate the playback order of the target video in the video collection; According to the target text, a key frame is determined in the target video, where the key frame is the target frame associated with the target text.
17. An electronic device, characterized in that: The system comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of the video processing method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the video processing method according to any one of claims 1 to 8.
19. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the video processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Video processing method and device
CN111757175A
Video identification method and device, computer equipment and storage medium
CN114782879A