Video key point extracting and playing method and device, medium and program product

By analyzing user playback behavior, identifying and extracting key plot segments from videos, the problem of users having difficulty locating key plot points is solved, enabling efficient skipping and playback of key plot segments and improving the user experience.

CN120935414APending Publication Date: 2025-11-11BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511050226.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Users often struggle to pinpoint exciting plot points when watching long videos, leading to wasted time and a diminished user experience.

Method used

By analyzing user playback behavior, especially playback speed adjustments, the system identifies and extracts exciting plot segments from videos and enables users to jump to these segments during playback.

Benefits of technology

It improves the efficiency of users watching exciting storylines, meets users' demand for exciting content, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935414A_ABST
    Figure CN120935414A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video key point extracting and playing method and device, a medium and a program product. The method comprises the following steps: acquiring a playing speed corresponding to each video clip in a target video; determining a switching video clip corresponding to a target switching behavior based on the playing speed, and determining a plurality of wonderful video clips and wonderful plot information; and respectively establishing association between the multiple pieces of wonderful plot information and the target video so as to jump from the target video to a time point corresponding to the wonderful plot segment and play when the target video is played on a playing page. According to the method and the device, the interested lens drama of the user can be determined by integrating the target switching behaviors of the plurality of users, and skip playing is performed among wonderful drama, so that the watching demand of the user on the wonderful drama is met, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and in particular to a method for extracting key points from a video, a video playback method based on key points, a video playback system, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the rise of short videos, users have gradually become accustomed to the short, fast, and concise pace of content, and are enthusiastic about the high-energy content continuously delivered by short video platforms. In contrast, traditional long videos, due to the needs of plot development, often contain a large amount of exposition and transitional content. This linear narrative mode is clearly out of step with the viewing habits of today's users. When users encounter segments with slow pacing or unengaging content, they easily lose patience and often skip uninteresting content by fast-forwarding or dragging the progress bar.

[0003] However, methods such as fast-forwarding and dragging the progress bar only adjust the playback time of the video, and the content that is adjusted is unknown. Therefore, users need to try repeatedly to find the content they want to watch. Blindly fast-forwarding can also easily cause them to miss important plot points and make it impossible to accurately locate exciting content, thus wasting time.

[0004] Therefore, a technical problem that urgently needs to be solved by those skilled in the art is: how to identify exciting plot points in a video so that playback can jump between these exciting plot points. Summary of the Invention

[0005] One objective of this invention is to provide a method for extracting key points from videos, accurately identifying key plot points within the video, enabling seamless playback transitions between these key plot points, and enhancing user interest and experience. The specific technical solution is as follows:

[0006] In a first aspect of this invention, a method for extracting key points from a video is provided, comprising: obtaining the playback speed corresponding to each video segment in a target video; determining a switching video segment corresponding to a target switching behavior based on the playback speed; determining multiple exciting video segments and exciting plot information based on the switching video segments corresponding to the target switching behavior, wherein the exciting plot information includes the time point corresponding to the exciting video segment; and establishing associations between the multiple exciting plot information and the target video, so that when the target video is played on the playback page, in response to a first trigger operation, the exciting plot segment closest to the current time point is determined based on the exciting plot information, and the playback is initiated from the target video to the time point corresponding to the exciting plot segment.

[0007] In a second aspect of the present invention, a video playback method based on key points is also provided, comprising: playing a target video on a playback page, the target video being associated with exciting plot information, the exciting plot information including: time points corresponding to exciting plot segments, the exciting plot segments being determined based on audio-visual content adjustment processing of exciting video segments, the exciting video segments being determined based on switching video segments corresponding to target switching behavior, and the target switching behavior being determined based on the playback speed corresponding to each video segment in the target video.

[0008] In a third aspect of this invention, a video playback system is also provided, comprising: a server and a client; wherein, the server acquires the playback speed corresponding to each video segment in a target video; determines the switching video segment corresponding to a target switching action based on the playback speed, and determines multiple exciting video segments and exciting plot information based on the switching video segments corresponding to the target switching action, the exciting plot information including the time point corresponding to the exciting video segment; and associates the multiple exciting plot information with the target video respectively; the client plays the target video on a playback page, and in response to a first trigger operation, determines the exciting plot segment closest to the current time point based on the exciting plot information; and jumps the target video to the time point corresponding to the exciting plot segment and plays it.

[0009] In another aspect of the present invention, an electronic device is also provided, characterized in that it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; the processor is used to execute the program stored in the memory to implement the steps of the video key point extraction method as described in the embodiments of the present invention, and to implement the steps of the video playback method based on key points as described in the embodiments of the present invention.

[0010] In another aspect of the present invention, a computer-readable storage medium is also provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform any of the above-described video key point extraction methods and key point-based video playback methods.

[0011] In another aspect of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the above-described video key point extraction methods and key point-based video playback methods.

[0012] The video key point extraction method provided in this invention obtains the playback speed corresponding to each video segment in the target video. Playback speed reflects the user's playback behavior; therefore, the switching video segment corresponding to the target switching behavior can be determined based on the playback speed. By combining the target switching behaviors of multiple users, multiple exciting video segments and plot information in the target video are determined, and the exciting plot points of interest to the user are identified from the user's switching behavior. Each piece of exciting plot information is associated with the target video so that when the target video is played on the playback page, in response to a first trigger operation, the playback jumps from the target video to the corresponding time point of the exciting plot segment and plays it. This allows for jumping between exciting plot points, satisfying the user's viewing needs for exciting plots and improving the user experience. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0014] Figure 1 This is a flowchart illustrating the steps of an embodiment of the video key point extraction method of the present invention;

[0015] Figure 2 This is a flowchart of a sub-step in determining a switched video segment based on target switching behavior in an embodiment of the present invention;

[0016] Figure 3 This is a schematic diagram illustrating the extraction of switched video segments in an embodiment of the present invention;

[0017] Figure 4 This is a flowchart of a sub-step in an embodiment of the present invention for determining a switching video segment based on target switching behavior;

[0018] Figure 5 This is a schematic diagram illustrating the overlapping of multi-user handover times in an embodiment of the present invention;

[0019] Figure 6 This is a bar chart of heat distribution in one embodiment of the present invention;

[0020] Figure 7 This is a schematic diagram illustrating an example of a heat curve in an embodiment of the present invention;

[0021] Figure 8 This is a flowchart of a sub-step in determining a highlight video clip according to an embodiment of the present invention;

[0022] Figure 9 This is a flowchart of the sub-steps of an adjustment process for a highlight video clip in an embodiment of the present invention;

[0023] Figure 10This is a flowchart of an optional embodiment of a video playback method based on key points according to the present invention;

[0024] Figure 11 This is a flowchart of a sub-step in another embodiment of the present invention for adjusting a highlight video clip;

[0025] Figure 12 This is a flowchart of an optional embodiment of another key-point-based video playback method of the present invention;

[0026] Figure 13 This is a flowchart of the secondary extraction step of exciting plot in an embodiment of the present invention;

[0027] Figure 14 This is a flowchart illustrating the steps of an embodiment of a video playback method based on key points according to the present invention.

[0028] Figure 15 This is a schematic diagram illustrating a plot jump example according to an embodiment of the present invention;

[0029] Figure 16 This is a schematic diagram illustrating an example of a playback page according to an embodiment of the present invention;

[0030] Figure 17 This is a flowchart illustrating the steps of another embodiment of the video playback method based on key points according to the present invention;

[0031] Figure 18 This is a schematic diagram of an embodiment of the video playback system of the present invention;

[0032] Figure 19 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0033] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0034] This invention provides a method for extracting key points from videos. By combining user playback behavior during video playback and analyzing actual user behavior characteristics, such as adjustments to playback speed, the method identifies key moments in the story that better reflect the user's actual preferences. These key moments are then associated with the target video. When the target video is played on the terminal, in response to a first trigger operation, the system jumps from the target video to the corresponding key moment in the story segment and plays it. This allows for seamless switching between key moments, satisfying the user's viewing needs and improving the user experience.

[0035] Reference Figure 1 The diagram illustrates a flowchart of an embodiment of a video key point extraction method according to the present invention.

[0036] Step 102: Obtain the playback speed of each video segment in the target video.

[0037] The target video to be analyzed is a video already played on a video playback platform. These platforms provide videos to users through video websites, applications, and apps. Therefore, the video playback platform stores playback information for this target video, which is determined based on each user's behavior while playing it. When watching the target video (through a website or app), users may fast-forward or skip segments they are not interested in, while for segments they are interested in, they usually adjust the playback speed to normal (or constant speed playback), or play those segments repeatedly. Some users may also play them slowly. Normal speed playback refers to playback at the normal speed, such as 1x speed. Fast playback refers to playback at a speed greater than normal, such as 1.25, 1.5, 2, 3, 4, or 5x speed. Slow playback refers to playback at a speed less than normal, such as 0.75 or 0.5x speed.

[0038] For a target video, multiple video segments are defined based on the playback speed at which each user plays the target video. Each video segment corresponds to a specific playback speed. For example, for a 45-minute (min) target video, if the first 3 minutes and the last 3 minutes are skipped, the corresponding playback speeds are: 0 for 3-6 minutes, 3x speed for 6-7 minutes, normal speed for 7-9 minutes, 2x speed for 36-40 minutes, and normal speed for 40-42 minutes. The video segments can be divided accordingly, such as: 0-3 min as segment 1 (0x speed), 3-6 min as segment 2 (3x speed), 6-7 min as segment 3 (normal speed or 1x speed), 7-9 min as segment 4 (2x speed), 36-40 min as segment (n-2) (3x speed), 40-42 min as segment (n-1) (normal speed), and 42-45 min as segment n (0x speed). That is, each playback speed adjustment point is a segmentation point, so each video segment corresponds to only one playback speed.

[0039] For each video segment, the playback speed and time range of that segment can be used as its playback information, thus obtaining the playback information for the target video. The video segments are arranged in chronological order within the playback information.

[0040] Step 104: Determine the switching video segment corresponding to the target switching behavior based on the playback speed; determine multiple exciting video segments and exciting plot information based on the switching video segment corresponding to the target switching behavior.

[0041] The playback information shows that adjacent video segments have different playback speeds; that is, the start and end times of each video segment are the switching points between different playback speeds. Target switching behavior refers to actions that indicate a user's interest in the storyline. Therefore, target switching behavior can be the action of switching the playback speed from high to low, such as switching from fast playback to normal playback speed.

[0042] The system can determine target switching behavior based on playback speed, and then select a pre-defined video segment for switching based on this behavior. This pre-defined video segment is likely to be of interest to the user. It can also integrate playback information of target videos within the pre-defined time range, and filter out multiple engaging video segments from these segments based on user behavior. These engaging video segments are of interest to multiple users. Here, "engaging plot" refers to compelling storyline content within the video.

[0043] For each exciting video clip, you can identify key plot information by adding the start time of the exciting video clip to the plot information, so that you can locate that exciting video clip.

[0044] Step 106: Associate multiple pieces of exciting plot information with the target video respectively, so that when the target video is played on the playback page, in response to the first trigger operation, the exciting plot segment closest to the current time point is determined based on the exciting plot information, and the playback is initiated from the target video to the time point corresponding to the exciting plot segment.

[0045] After extracting key plot information, this information can be associated with the target video. For example, jump points can be set in the target video based on the time points of the key plot segments in the key plot information. Alternatively, the key plot information can be bound to a pre-defined action, allowing users to jump between different time points based on that action. Thus, during video playback, upon receiving a pre-defined action, the system can jump to the next time point and begin playing video data from that point. This allows users to skip to the next key plot point when they are no longer interested in the current one, satisfying their viewing needs and improving the user experience.

[0046] In summary, by obtaining the playback speed of each video segment in the target video, and recognizing that playback speed reflects user behavior, the corresponding video segment for the target switching behavior can be determined based on the playback speed. By combining the target switching behaviors of multiple users, multiple exciting video segments and plot information within the target video can be identified. From the user's switching behavior, the exciting plot points that the user is interested in can be determined. These plot points are then associated with the target video so that when the target video is played on the playback page, in response to a first trigger operation, the user can jump from the target video to the corresponding time point of the exciting plot segment and start playing it. This allows for seamless switching between exciting plot points, satisfying the user's viewing needs and improving the user experience.

[0047] This invention analyzes video segments of interest based on user behavior. When users are browsing video content at high speed, they often actively slow down the playback speed to watch the segments they are interested in more closely. This behavior pattern of switching from high speed to normal or slow speed actually reflects the user's subjective evaluation of the content.

[0048] In one optional embodiment, step 104, determining the switching video segment corresponding to the target switching behavior based on the playback speed, includes:

[0049] Sub-step 201: Determine the time point corresponding to the target switching behavior, and obtain a switching video segment of predetermined duration starting from the time point. The target switching behavior includes: a first behavior of switching from fast playback to normal or slow playback, and / or a second behavior of switching from normal playback to slow playback.

[0050] Therefore, a first time point can be determined for switching from fast playback to normal or slow playback, and a switching video segment of predetermined duration can be obtained starting from the first time point; and / or, a second time point can be determined for switching from normal playback to slow playback, and a switching video segment of predetermined duration can be obtained starting from the second time point. The playback speeds of two adjacent video segments can be determined. When the playback speed of the preceding video segment is less than the playback speed of the following video segment, the end time point of the preceding video segment or the start time point of the following video segment can be determined as the time point corresponding to the target switching behavior.

[0051] By using the user's behavior of switching from high speed to normal speed as a marker of content attractiveness, this method can accurately identify engaging content directly from the user's natural interaction behavior without the need for complex content understanding algorithms. The method is simple and reliable.

[0052] In this context, the video clip transition is referred to as a "speed bump." For example, a pre-defined duration of video clip transitioning from 3x speed playback to normal speed playback is called a "speed bump." The pre-defined duration is the length of time sufficient to cover the entire story segment, such as 20, 25, 30, 35, or 40 seconds. This example uses a transition from 3x speed playback to normal speed playback; in actual processing, this could involve adjusting the playback speed from high to low, without limiting the specific playback speed value.

[0053] by Figure 3 For example, the first timeline represents playback speed, with 3x speed represented in black and normal speed in white. The second timeline represents the timing information of video segments that the user is interested in, with diagonal lines representing normal video segments and vertical lines representing switched video segments.

[0054] A user's voluntary slowing down of playback speed indicates their interest in a video segment. Based on this interest, a time window of a set duration can be set. Thus, the video segment to switch to is determined based on the timing of the playback speed reduction and the set time window.

[0055] In this embodiment of the invention, it is also possible to detect whether there is a switching point in the predetermined duration video segment where the playback speed is increased. If such a switching point exists, the video segment is ignored. That is, if the user speeds up again within the predetermined duration, it indicates that the plot content lacks sustained appeal and the segment needs to be terminated early, and is not considered a video segment of interest.

[0056] In another optional embodiment, the target switching behavior further includes a rewind behavior, i.e., a third behavior of switching back from the later video segment to the previous video segment. This rewind behavior adjusts the playback time from the end to the beginning. After adjusting the playback time from the end to the beginning, the user may also adjust the playback speed, so it may be a combination of multiple switching behaviors. The playback information also includes rewind information corresponding to the rewind behavior, which includes a start time point, at which a pre-defined duration of the switched video segment can be obtained. Since the user may also adjust the playback speed after the rewind behavior is executed, multiple switching behaviors may also be combined. The switched video segment can be determined by combining multiple behaviors.

[0057] After determining the video segment to switch based on each user's playback behavior, the behavior of multiple users can be combined to identify several exciting video segments within the target video, thus providing users with corresponding engaging storylines. Therefore, playback data of the target video within a set time period can be obtained, and the playback behavior of each user within that time period can be analyzed, integrated, and statistically analyzed to determine the exciting video segments.

[0058] In one alternative embodiment, such as Figure 2 As shown, step 104, based on the target switching behavior corresponding to the switching video segment, determines multiple exciting video segments, including:

[0059] Sub-step 203: The target switching behavior of each user is statistically analyzed in chronological order to determine the playback popularity curve of the target video;

[0060] Sub-step 205: Based on the playback popularity curve, determine multiple exciting video clips.

[0061] The target video switching behaviors of each user within a set time period are sorted according to the chronological order of the corresponding time points of the target switching behaviors. That is, the switched video segments are sorted according to the order of their start times. Each user can be assigned a weight, and each time point is weighted according to the user's weight to obtain the weight information for that time point. Based on the time point weight information, a playback popularity curve for the target video is determined. This popularity curve is used to represent the level of interest of the video's plot content at different time points. For example, different weights can be assigned based on different user levels. Then, multiple time points are filtered according to this playback popularity curve, and the switched video segments corresponding to those time points are selected as the most exciting video segments.

[0062] In another alternative embodiment, such as Figure 4 As shown, in step 104, based on the target switching behavior corresponding to the switching video segment, multiple exciting video segments are determined, including:

[0063] Step 207: Based on the video segments switched by each user, count the number of overlaps at time points, and generate the playback popularity curve of the target video based on the number of overlaps;

[0064] Sub-step 209: Based on the playback popularity curve, determine multiple exciting video clips.

[0065] Sub-step 211: Obtain the start time of the exciting video clip as information for the exciting plot.

[0066] The time information corresponding to the switched video segments for each user is determined. This time information is the time information between the start time point and the end time point, such as the timestamp information between the start time point and the end time point. Then, the number of switched video segments corresponding to each time point is determined as the number of overlaps of the time points, which is used as the popularity value. The playback popularity curve of the target video is determined according to the popularity value of the time points.

[0067] In this embodiment of the invention, by statistically analyzing the deceleration behavior of a large number of users in the same video, a popularity curve reflecting the excitement of the plot can be constructed. In one example, the video clips corresponding to the deceleration behaviors of multiple users are overlaid and statistically analyzed over a time dimension, such as... Figure 5 As shown, User 1 exhibited deceleration behavior in multiple time intervals, including 0-20s, 30-60s, and 80-110s; User 2 exhibited deceleration behavior in multiple time intervals, including 20-30s, 60-80s, and 100-120s; and User 3 exhibited deceleration behavior in multiple time intervals, including 10-30s, 50-80s, and 90-110s. The deceleration patterns of different users overlapped on the timeline.

[0068] Then, the number of overlaps per second between different users' speed bumps can be counted, which is the popularity value, resulting in a popularity distribution map reflecting the content's appeal. Figure 6 The image shows a bar chart illustrating the popularity distribution. In this chart, the horizontal axis represents time, and the vertical axis represents popularity values, such as those based on overlap counts or weights. This bar chart reflects user switching patterns at different time points, thus reflecting the popularity of the video segment at the corresponding time point. Popularity distribution charts can also be displayed as time-point distribution charts, and this embodiment of the invention does not impose limitations on this.

[0069] This invention introduces the concept of speed bumps and constructs a popularity distribution by superimposing the time dimension of multiple user behaviors. This effectively filters out the random behavior of individual users, highlights truly exciting content through the collective wisdom of users, and improves the accuracy of identification.

[0070] When the time points are sufficiently close together, a continuous popularity curve can be formed. For example, a 45-minute episode of a TV series has a playback time of approximately 2700 seconds. Figure 7 In the heat curve diagram shown, the horizontal axis is the time axis, and the vertical axis is the number of times playback was paused at double speed. Pausing playback at double speed can be understood as a target switching behavior, with the vertical axis representing the number of target switching behaviors. The orange curve is generated based on a specific test episode. The peak of this curve indicates that multiple users slowed down their viewing at similar times. The green curve (PDPM, or Playback Duration Per Minute) is also higher at higher points on the orange curve, indirectly confirming the strong appeal of the content at that point. The PDPM curve is formed by combining various user playback actions such as normal / double speed playback, video dragging, fast forwarding, and rewinding. It reflects the user's viewing time at the minute level of the progress bar; the higher the value, the more interested the user is in the content. However, the granularity (i.e., precision) of the PDPM curve is in minutes, which is relatively low and insufficient to pinpoint the timestamps of exciting plot points at the second level. This embodiment of the invention uses timestamps at specific times as the precision, accurate to two decimal places, resulting in higher precision.

[0071] After determining the popularity curve, multiple highlight video segments can be identified based on it. For example, multiple intervals can be defined, and the video segment corresponding to the highest point in the popularity curve within each interval can be extracted as the highlight video segment. Alternatively, target time points can be extracted by combining mean and peak values, and then the video segment containing the target time can be determined as the highlight video segment.

[0072] In one alternative embodiment, such as Figure 8 As shown, sub-step 1045 or sub-step 1049 determines multiple highlight video clips based on the playback popularity curve, including:

[0073] Sub-step 802: Determine the average value of the playback popularity corresponding to the playback popularity curve.

[0074] The average value can be determined based on the hot points in the playback popularity curve. For example, hot points can be characterized by the number of overlaps or the number of playback switches. The average playback popularity can be calculated based on the hot points.

[0075] In this embodiment of the invention, the effective content of the target video is pre-divided, including deleting the intro and outro portions of the target video. The intro and outro portions typically contain non-plot content such as the title and cast list, which are not highly relevant to the plot and can therefore be ignored. For example, the first 3-4 minutes and the last 3-4 minutes—that is, the intro video segment starting from the beginning of the target video and the outro video segment starting 3-4 minutes before the end of the target video.

[0076] Sub-step 804: Based on the average value, determine the candidate time series of the popularity peak corresponding to the playback popularity curve. The candidate time series includes the timestamps of candidate time points where the popularity peak is higher than the average value.

[0077] After determining the average value, we can identify the data points in the popularity curve that are above the average value. Then, we can determine the candidate time points where the popularity peaks occur from these data points, and finally, we can determine the candidate time series based on the subsequent time points. The popularity peak refers to the highest point within a certain interval of the popularity curve, which can be divided into multiple intervals based on time, popularity, etc.

[0078] In one optional embodiment, multiple peaks can be determined in the playback popularity curve based on a peak detection algorithm. Each peak is compared with the average value, and the peaks that are greater than the average value are determined as popularity peaks. The time points corresponding to the popularity peaks are used as candidate time points to obtain candidate time series.

[0079] Sub-step 806: Filter the candidate time series according to a predetermined time window to determine multiple target time points.

[0080] A predetermined time window can be set, the size of which is related to the predetermined duration. For example, the size of the predetermined time window is 1.5 minutes, 2 times, or more of the predetermined duration. The candidate time series are filtered according to the predetermined time window. For example, after sorting the candidate time series in chronological order, they are filtered according to the predetermined time window, and the time point with the highest popularity value within each predetermined time window is selected as the target time point.

[0081] In another optional embodiment, candidate time series are sorted from highest to lowest according to the popularity value corresponding to each time point, and then filtered according to a predetermined time window. The predetermined time window is set with the median point of the specified time point as its position, and then the time point corresponding to the highest popularity value is selected as the target time point within the predetermined time window. In one example, the time points corresponding to the peak popularity are sorted from highest to lowest popularity value, and processing begins from the time point with the highest popularity value. For each time point to be processed, taking a predetermined duration of 30 seconds as an example, it is checked whether there are other candidate points within the predetermined time window of 30 seconds before and after it (a total of 60 seconds). If so, these candidate points with lower popularity values ​​are suppressed, that is, the candidate points with lower popularity values ​​are deleted from the candidate set, and only the time point with the highest current popularity value is retained. This processing method ensures that at most one most representative time point, i.e., the highlight point, is retained within any 60-second time window, thereby effectively avoiding the problem of overly dense point distribution.

[0082] Sub-step 808: Obtain multiple exciting video clips corresponding to the multiple target time points respectively.

[0083] Starting from the target time point, obtain the video segment corresponding to the predetermined duration from the starting time point, and use the video segment of the predetermined duration as the highlight video segment to obtain multiple highlight video segments.

[0084] Since the plot point corresponding to a user's switching behavior is uncertain—it could be in the middle of a storyline or a scene with an empty shot—it's possible to adjust the audio-visual content of exciting video clips to ensure a complete story segment is provided.

[0085] While user playback behavior can reflect their interest in the video's storyline, user switching behavior often makes it difficult to pinpoint the start and end of the story. Switches might occur midway through a scene, such as during a dialogue, or in an empty area, such as during a scene transition. Therefore, relying solely on user behavior to determine key storylines and presenting them directly to other users could potentially confuse them and degrade the user experience.

[0086] Therefore, based on the above embodiments, this embodiment of the invention also detects based on the audio-visual content of the video clip itself. If it is found that the starting time point of an exciting plot segment is not accurately located, which will affect the user's viewing experience, the starting time point will be adjusted to a suitable position to improve the user experience.

[0087] In one optional embodiment, the plurality of exciting video clips are adjusted based on their audio-visual content to determine a plurality of exciting plot information, the exciting plot information including the time points corresponding to the exciting video clips.

[0088] The audio-visual content refers to the content presented in the video through visuals and audio. By combining the audio-visual content of a highlight video clip and adjusting the clip from multiple dimensions, including visual content and audio, the key plot points can be identified. By repositioning the start and end times of the video clips, the corresponding key plot points and time information can be obtained, improving the accuracy of key plot point positioning and thus enhancing the user experience.

[0089] In one optional embodiment, step 106 involves adjusting the plurality of exciting video clips based on their audio-visual content to determine multiple exciting plot details, including:

[0090] Step 1062, adjust the picture of the exciting video clip to determine the exciting plot information of the exciting plot clip starting from the plot scene; and / or, Step 1064, perform dialogue continuity detection on the exciting video clip to determine the exciting plot information of the exciting plot clip starting from the dialogue access point, wherein the dialogue access point includes: dialogue end point and / or dialogue start point.

[0091] Determine the starting time point corresponding to the exciting video clip, obtain the video image frame or video image frame description information corresponding to the timestamp of the starting time point, analyze the plot content corresponding to the starting time point based on the video image frame, and determine whether it is the scene where the plot content is located or an empty shot. If it is an empty shot, adjust the starting time point to the next timestamp that is not an empty shot. Use this timestamp as the starting time point to determine the exciting video clip, which can also be called an exciting plot clip, and obtain the time information and other exciting plot information of the exciting plot clip.

[0092] The system acquires audio data corresponding to key video clips. For example, it extracts audio data within a specified duration (e.g., 10 seconds) starting from the beginning of the key video clip, or extracts audio data before and after a specified duration from the beginning of the key video clip. Based on the audio data, it identifies text information and detects dialogue continuity, such as whether important dialogue or plot turning points have been interrupted. If there is a risk of interrupting dialogue, the timestamp of the switching point is adjusted, and this timestamp is used as the starting point to determine the key video clip, also known as a key plot segment. This key plot segment includes its time information and other plot details.

[0093] It can also combine screen adjustments and dialogue continuity detection to identify exciting plot segments and information.

[0094] Based on the above embodiments, this invention provides an optional embodiment of a video key point extraction method, referring to... Figure 10 As shown:

[0095] Step 1002: Obtain the playback speed corresponding to each video segment in the target video.

[0096] The target video has corresponding playback information, which includes video segment information divided according to playback speed, including double speed playback and normal speed playback.

[0097] Step 1004: Determine the time point corresponding to the target switching behavior, and obtain a switching video segment of a predetermined duration starting from the time point.

[0098] Step 1006: Based on the video segments switched by each user, count the number of overlaps at time points, and determine the playback popularity curve of the target video based on the number of overlaps.

[0099] Step 1008: Determine multiple exciting video clips based on the playback popularity curve.

[0100] Step 1010: Perform image detection on the exciting video clips to determine the exciting plot information of the exciting plot clips starting from the plot scenes.

[0101] Step 1012: Perform dialogue coherence detection on the exciting video clips to determine the exciting plot information of the exciting plot clips starting from the dialogue access point.

[0102] In one optional embodiment of the present invention, video description information corresponding to a key video clip is determined. This video description information describes the plot content of the video clip. The video description information can be obtained by identifying the video data corresponding to the video clip, including the identification of video images and audio data. Alternatively, the video description information can be determined in conjunction with plot information provided by the producer of the target video; this embodiment of the present invention does not impose limitations on this approach.

[0103] The description of the plot content in a video clip can be divided into multiple plot units. For example, the video clip can be divided into plot units every 15-20 seconds. Each plot unit describes the plot from multiple dimensions such as visuals and audio, thus refining the video clip into individual plot units. This detailed information organization provides a data foundation for subsequent analysis, improving the accuracy and comprehensiveness of the analysis. Each plot unit corresponds to time information and audio-visual content. The time information of a plot unit can be the time range of the plot or a specific time point. The time point can be the start time of the plot or other specified time points within the plot, such as the time points containing characters. In an optional embodiment, the time point of the plot unit is a timestamp. The plot units can be sorted according to their time points in the video description information, that is, the video description information includes a sequence of time points for plot units, with each time point in the sequence corresponding to a plot unit, and each plot unit is described through audio-visual content.

[0104] Based on video description information, the semantic understanding of audio and video content is used to determine the plot content presented. By detecting the scene and dialogue continuity of exciting video clips, exciting plot segments and their exciting plot information can be identified.

[0105] In one optional embodiment, step 106 involves adjusting the multiple exciting video clips based on their audio-visual content to determine multiple exciting plot information, including the following sub-steps: performing multi-dimensional plot understanding on the exciting video clips based on a preset language model, and outputting the adjusted exciting plot information.

[0106] This invention employs a pre-trained language model to detect plot information such as video content and dialogue coherence. The language model is an abstract mathematical model constructed based on objective linguistic facts and is a fundamental tool in computational linguistics. Natural Language Processing (NLP) is performed based on the language model to understand the semantics of the video description information, thereby extracting key plot information from the video clips. The pre-trained language model can be various NLP-based language models, such as Transformer-based NLP models, multi-task learning models, neural network models, and Large Language Models (LLMs). The LLM model is a deep learning model trained using large amounts of text data, enabling it to generate natural language text or understand the meaning of linguistic text.

[0107] Input information is constructed based on video description information and fed into a pre-built language model for plot understanding processing. The model analyzes key video clips from multiple dimensions, including starting screen information and dialogue coherence, and synthesizes the results to determine key plot segments and related information. Furthermore, by analyzing these key plot segments using the pre-built model, titles can be determined. These titles are summaries of the audio-visual content of the key plot segments, such as a summary of the plot in no more than 10-20 words.

[0108] Optionally, the step of performing multi-dimensional plot understanding on the exciting video clips based on a pre-built language model and outputting the adjusted exciting plot information includes:

[0109] Sub-step 1102: Obtain video description information corresponding to the exciting video clip, the video description information including audio and video content corresponding to the exciting video clip.

[0110] Sub-step 1104: Obtain the prompt word template. The prompt word template includes variable parameters and prompt items. The prompt items include at least one of the following: analysis content, screening criteria, plot screening principles, and output requirements.

[0111] Pre-built language models can be combined with prompt engineering to perform analysis and processing. Prompt engineering is used to design and adjust input information, guiding pre-built language models such as LLM models to generate the desired output. Prompt templates can be designed using prompts, and these templates can then guide pre-built language models such as LLM models.

[0112] This prompt template can serve as the primary prompt template. The primary prompt template includes variable parameters and prompt items. Variable parameters are the variables that need to be filled into the prompt template. The settings of these variables are related to the video description information. The variable parameters are associated with the prompt items, and the analysis and processing operations for the video description information are determined based on the prompt items. For example, a time variable is used to fill in the switching time point, such as the start time point of a video clip. Alternatively, the audio-visual content corresponding to the start time point can be determined and filled into the content variable, such as filling the dialogue content into the dialogue variable.

[0113] The prompts include at least one of the following: analysis content, evaluation priority, and output requirements. The analysis content guides the pre-built language model in performing the required analysis and processing operations. The analysis content includes analysis prompts and target information. The analysis prompts indicate the items and content to be analyzed, while the target information indicates the objectives the analysis must meet. The evaluation priority is the priority for evaluating the plot content of the exciting video clips. For example, the priority from highest to lowest is: event causal logic information, scene coherence, dialogue integrity, and other factors. Event causal logic information determines the integrity and causal logic of the event, generally requiring conditions such as including the cause and development of the event and maintaining causal relationships. Scene coherence determines the degree of scene coherence, such as prioritizing the start time of a new scene or the time of scene transitions. Dialogue integrity determines the completeness of the dialogue content corresponding to the starting time, avoiding interruptions in dialogue or switching between speakers' utterances. Other factors are other factors affecting the viewing experience, such as whether important lines are included.

[0114] In one example, the variable parameters in the prompt template are marked with {{}}, and the values ​​inside the curly braces are to be filled in. For example, variables related to the current point in time, such as the dialogue content and screen information corresponding to the start time, are to be filled in, helping the model understand upcoming and skipped plot information. An example of a prompt template for detecting dialogue coherence is described below:

[0115] You are a dialogue analysis expert. To help users implement the "Quick View" feature, please analyze the dialogue content before and after the current starting point ({{current_point}}), and find the most suitable starting point from the existing dialogue timestamps. A good starting point should allow users to fully understand a short story / scene after jumping to the next point. If the current starting point is not suitable, please find a more appropriate entry point in the dialogue.

[0116] Core objective:

[0117] - Select the best jump point from the existing conversation timestamps.

[0118] - Ensure that viewers can quickly understand the scene and events after switching to another page.

[0119] It's better to include more context than to confuse the audience.

[0120] Explanation of start time point:

[0121] - The starting time point ({{current_point}}) is the starting point from which the viewer actually begins watching each jump.

[0122] - Content preceding the starting time point is invisible to the viewer and may contain preceding text from the starting time point.

[0123] -Assess whether it is easy to understand from the perspective of a first-time viewer.

[0124] - Only timestamps already existing in the conversation can be selected; creating or imagining new timestamps is prohibited.

[0125] Main character list: {{charac_names}}

[0126] The dialogue at the starting time point ({{current_point}}: {{cur_line}}

[0127] Evaluation priority (from high to low):

[0128] 1. Event completeness and causal relationship

[0129] - The cause and development of the event must be included.

[0130] - Cannot be approached from the event process

[0131] -Contents with causal relationships cannot be separated.

[0132] - Ensure the audience understands "why this happened".

[0133] 2. Scene coherence

[0134] - Prioritize starting from scene / environment changes

[0135] - The dialogue at the beginning of a new scene is more suitable

[0136] - Avoid starting from the middle of the scene

[0137] 3. Dialogue integrity

[0138] - Avoid cutting off question-and-answer pairs

[0139] - Avoid interrupting the same speaker's continuous speech.

[0140] - Avoid disrupting the semantic flow of the conversation

[0141] 4. Other considerations

[0142] -Important lines from main characters

[0143] -A natural transition point in the topic

[0144] - Opening remarks such as titles and greetings

[0145] Special reminder:

[0146] Even if you find some great lines from the main characters, you can't sacrifice the integrity of the story for that.

[0147] It's better to include more context than to confuse the audience.

[0148] -The first question viewers will have after switching channels is "What's going on?", not "That line was well-said."

[0149] === ...

[0150]

[0151]

[0152] === ...

[0153] {{invisible_dialogues}}

[0154] === ...

[0155] {{visible_dialogues}}

[0156] === ...

[0157] {{clips_data}}

[0158] ======Other Notes=====

[0159] The returned timestamp can only come from the dialogue timestamp or the original starting time point; the start time of the event cannot be inferred arbitrarily.

[0160] The aforementioned prompt template can be obtained and subsequently combined with relevant video clips for further adjustments. This embodiment of the invention can guide the processing of a pre-built language model by setting prompt word templates. The prompt word templates can be flexibly set based on analysis needs. This guides the pre-built language model to perform accurate analysis and processing, improving the accuracy of the processing results.

[0161] Sub-step 1106: Combine the video description information and the prompt word template to generate input information.

[0162] The video description information and prompt word templates of exciting video clips are combined. The video description information is filled into the corresponding face-changing parameters to generate the input information of the pre-set language model.

[0163] Sub-step 1108: Input the input information into a preset language model for multi-dimensional plot understanding, and output adjusted exciting plot information, which includes: time information and title of exciting plot segments.

[0164] The input information, constructed from prompt word templates and video description information, is fed into a pre-built language model for analysis. The pre-built language model can combine the guidance of prompt words to perform multi-dimensional analysis and processing of the plot scenes and dialogue content corresponding to the starting time point of the exciting video clip, thereby determining the adjustment starting time point. Based on the adjustment starting time point, the exciting plot clip and time information are determined. It can also summarize the title of the exciting plot clip, thereby generating exciting plot information based on the time information and title.

[0165] In one example, dialogue content within 10 seconds before and after the start time of a key video clip can be extracted and analyzed using a large model. The model determines whether the current start time would disrupt the dialogue, such as cutting off important dialogue or plot twists. If a risk of dialogue interruption is detected, the model suggests a better entry point, typically where the dialogue ends naturally or a new scene begins. This post-processing ensures not only visual coherence but also narrative integrity. The adjustment effect of one example is shown in Table 1 below:

[0166]

[0167]

[0168] Table 1

[0169] This allows us to analyze the information in the scenes and dialogues corresponding to exciting video clips based on the model, select appropriate starting time points, and thus determine the exciting plot information of exciting story segments.

[0170] By using large-scale model analysis to examine the context of dialogues and adjusting the start times of exciting plot segments, the integrity of these segments in terms of plot and dialogue is ensured, thus improving the smoothness of the viewing experience.

[0171] Based on the above embodiments, this invention provides another optional embodiment of a video key point extraction method, as described above. Figure 12 As shown:

[0172] Step 1202: Obtain the playback speed of each video segment in the target video.

[0173] The playback information includes video segment information divided according to playback speed, wherein the playback speed includes double speed playback and normal speed playback.

[0174] Step 1204: Determine the time point corresponding to the target switching behavior, and obtain a switching video segment of a predetermined duration starting from the time point.

[0175] Step 1206: Based on the video segments switched by each user, count the number of overlaps at time points, and determine the playback popularity curve of the target video based on the number of overlaps.

[0176] Step 1208: Determine multiple exciting video clips based on the playback popularity curve.

[0177] Step 1210: Obtain the video description information corresponding to the exciting video clip, the video description information including the audio and video content corresponding to the exciting video clip.

[0178] Step 1212: Obtain the prompt word template. The prompt word template includes variable parameters and prompt items. The prompt items include at least one of the following: analysis content, evaluation priority, and output requirements.

[0179] Step 1214: Combine the video description information and the prompt word template to generate input information.

[0180] Step 1216: Input the input information into a preset language model for multi-dimensional plot understanding, and output adjusted plot information, which includes: time information and title of the plot segment.

[0181] After identifying key plot segments, if these segments are too sparse or too dense, it may negatively impact the user's viewing experience. When users jump between these segments, a sparse distribution might cause them to miss certain scenes, while excessive density could make it difficult to distinguish between key and ordinary scenes. Therefore, in one optional embodiment of this invention, a set time range can be established between different key segments. The maximum time point within this range is a first time threshold, and the minimum time point is a second time threshold. Using the start time of each key plot segment as a base time point, the start times of each segment can be sorted chronologically, and the time interval between adjacent time points can be calculated. If the time interval exceeds the first time threshold, the two corresponding key plot segments are considered sparse segments; if the time interval does not exceed the second time threshold, they are considered dense segments. Both situations require adjustment.

[0182] For two adjacent target time points with a time interval exceeding the first time threshold, their corresponding two key plot segments are sparse segments. The video segment to be detected between these two adjacent target time points can be determined, and secondary extraction of key plot information can be performed on this segment. For two adjacent target time points with a time interval not exceeding the second threshold, their corresponding two key plot segments are dense segments. These two dense segments can be merged, and the key plot segments and their information can be re-determined. This can be achieved through the following steps:

[0183] Step 1302: Arrange the start times of the exciting plot segments extracted from each video clip in chronological order.

[0184] After selecting multiple exciting video clips, they can be sorted according to their starting time to generate a sequence of starting time points for the exciting video clips.

[0185] Step 1304: Determine the time interval between two adjacent starting time points.

[0186] Step 1306: Determine whether the time interval is within the set time range.

[0187] If the time interval is greater than the maximum time point of the set time range, i.e., the time interval is greater than the first time threshold, proceed to step 1308. If the time interval is less than the minimum time point of the set time range, i.e., the time interval is less than the second time threshold, proceed to step 1312. If the time interval is within the set time range, return to step 1304 to continue determining the time interval between the next two adjacent time points.

[0188] Step 1308: Determine the video description information of the video segment to be detected between the two adjacent target time points.

[0189] The video segment to be detected is determined between two adjacent target time points, and video description information of the video segment to be detected is obtained. The video description information includes the audio and video content of the video segment to be detected.

[0190] Step 1310: Based on the audio-visual content, perform multi-dimensional plot understanding on the video segment to be detected to determine exciting plot information.

[0191] Step 1312: Merge the video segments between the two adjacent target time points to generate the video segment to be detected.

[0192] The video segments between the two adjacent target time points are merged, and the merged video segments are then re-filtered to generate a video segment to be detected. The video description information of this video segment to be detected is obtained. The video description information includes the audio and video content of the video segment to be detected.

[0193] Step 1314: Based on the audio-visual content, perform multi-dimensional plot understanding on the video segment to be detected to determine exciting plot information.

[0194] The audio-visual content of the video description information for each video segment to be tested is analyzed separately. Semantic understanding of the audio-visual content is used to determine the presented plot content. Then, the overall plot content of the video segments to be tested is evaluated to determine the role of each plot point in driving the main plot development, thereby selecting the most compelling plot segments containing the climax. These compelling plot segments refer to the most attractive parts of the plot. Based on the time points and audio-visual content of these compelling plot segments, exciting plot information is generated.

[0195] In one optional embodiment, based on the audio-visual content, a multi-dimensional narrative understanding is performed on the video segment to be detected to determine key narrative information, including: performing a multi-dimensional narrative understanding on the video description information based on a pre-set language model, and outputting key narrative information. The step of performing a multi-dimensional narrative understanding on the audio-visual content based on a pre-set language model and outputting key narrative information corresponding to the video segment includes:

[0196] Obtain a second prompt word template, which includes variable parameters and prompt items. The prompt items include at least one of the following: analysis content, screening criteria, plot screening principles, and output requirements. Combine the video description information and the second prompt word template to generate input information. Input the input information into a preset language model for multi-dimensional plot understanding and output the exciting plot information corresponding to the video segment. The exciting plot information also includes the title of the exciting plot segment.

[0197] The prompts include at least one of the following: analysis content, screening criteria, plot selection principles, and output requirements. The analysis content guides the pre-built language model in performing the required analysis and processing operations. The analysis content may also describe the provided basic data, such as video description information. The screening criteria guide the pre-built language model in executing the required screening criteria for analysis and processing. The plot selection principles guide the pre-built language model in executing the required plot selection principles for analysis and processing. The plot selection principles can be set similarly to those described above, and include at least one of the following: integrity principle, independence principle, attractiveness principle, rhythm principle, and viewing experience principle. Specific details can be found in the description of the embodiments above, and will not be repeated here. The output requirements determine the content and format of the output results of the pre-built language model, such as requiring the model to input the time point, title, and selection reason for the selected exciting plot segments.

[0198] In one optional embodiment, the aforementioned second prompt template can be a universal prompt template for various drama series types, with the selection criteria for that drama series type being set by configuring type variables. This second prompt template is designed to meet the purpose and needs of video analysis, including guidance on identifying characteristics of different genres. For example, for fantasy dramas, it pays particular attention to elements such as emotional bonds and challenges faced by the protagonist, while for suspense dramas, it focuses more on clue development and plot twists. In another optional embodiment, a prompt template can be set for each drama series type, with the corresponding selection criteria set for that type, thereby directly obtaining the prompt template for the corresponding drama series type.

[0199] This invention can guide the processing of a pre-built language model by setting a second prompt word template. The prompt word template can be flexibly set based on analysis requirements. This guides the pre-built language model to perform accurate analysis and processing, improving the accuracy of the processing results.

[0200] This invention presents a multi-level time point filtering algorithm. It addresses the problem of densely packed points by using non-maximum suppression, and supplements it with a time point filtering method based on multi-dimensional model analysis to handle sparsely populated areas, ensuring that the interval between adjacent points is maintained at an optimal viewing rhythm.

[0201] In terms of operations, the aforementioned extraction technology can assist operators in quickly screening and editing video materials, improving work efficiency and providing material support for short video secondary creation. In one optional embodiment, the video data of the target video can also be edited based on the aforementioned compelling plot information to generate a video of a set duration. This compelling plot information can provide material support for editing; the plot in the target video can be screened based on this information and re-edited to obtain a video of a set duration, such as behind-the-scenes footage, plot introduction videos, and other secondary creation videos.

[0202] In one optional embodiment, targeting points can be determined based on the exciting plot information, and targeting information can be added to these targeting points. The plot type within the target video is determined based on the exciting plot information, such as a climax, a calmer part, or a transitional part, thereby setting corresponding targeting points. Targeting information, such as advertisements, is then set and played at these targeting points, ensuring advertising effectiveness without affecting user experience.

[0203] In this embodiment of the invention, a content quality quantification mechanism based on user speed-up behavior is established. The concept of "speed bumps" is introduced into long video scenarios to capture the behavioral characteristics of users switching from speed-up playback to normal speed playback and slow speed playback. A popularity curve is drawn, and the intelligent perception of user viewing hotspots is achieved by extracting local peaks. Moreover, the quality of the points will become higher and higher as user data accumulates.

[0204] This invention uses user behavior data as one of the filtering criteria, enabling dynamic adjustment of the identification of exciting plot segments. This allows for more accurate capture of group interests and provides a better viewing experience. It can also be extended to personalized scenarios, providing differentiated recommendations of exciting plot segments based on different users' viewing habits.

[0205] The embodiments of the present invention implement an intelligent content navigation mechanism, which will significantly improve the user's viewing experience and stickiness, provide technical support for video platforms, and provide strong support for enhancing product competitiveness.

[0206] Based on the above embodiments, this invention also provides a video playback method based on key points, which can jump between key plot points during video playback after extracting key plot information, providing users with a better viewing experience.

[0207] Reference Figure 14 The diagram illustrates a flowchart of an embodiment of a video playback method based on key points according to the present invention.

[0208] Step 1402: Play the target video on the playback page. The target video is associated with exciting plot information.

[0209] After associating the target video with the exciting plot information, if the server receives a playback request from the client, it can send the video data of the target video to the client, which will then parse and render the video data and play the target video on the playback page.

[0210] like Figure 15 As shown, the target video is associated with exciting plot information. For example, the exciting plot segment of plot 1 is located at 22 minutes and 30 seconds into the video, with the title "B angrily rebukes A". The exciting plot segment of plot 2 is located at 23 minutes and 54 seconds into the video, with the title "A discovers the spirit stone".

[0211] The target video is associated with exciting plot information. The method for extracting the exciting plot information is the same as the above embodiment and will not be repeated here.

[0212] When playing a target video, mobile phones, tablets, and other terminal devices can enable a skip function under a set model. In one optional embodiment of the present invention, when the terminal device is in landscape mode, at least one skip operation area is set in the playback page. The skip operation area is the operation area for triggering skipping to play exciting plot segments. The position of this area can be set according to needs, for example, it can be set on the left side, the lower right side, or the top and bottom sides of the playback page. Alternatively, when the terminal device switches to landscape mode, a prompt message can be displayed to remind the viewer to trigger skipping to play exciting plot segments.

[0213] In one example, operation prompts can be displayed on the playback page, such as... Figure 16 In the example shown, there are navigation areas on the left and right sides of the playback page, with prompts such as "Scan up and down to jump to different storylines." A brightness control area is located in the second area from the left, with the prompt "Slide up and down to adjust brightness," and a volume control area is located in the second area from the right, with the prompt "Slide up and down to adjust volume." This allows users to adjust brightness, volume, and jump between different storylines by performing actions in the corresponding areas while watching the video.

[0214] Step 1404: In response to the first trigger operation, determine the most recent exciting plot segment based on the exciting plot information.

[0215] While watching a target video on the playback page, if a user wants to switch to another scene, they can perform a first trigger action. The corresponding client can receive this first trigger action, such as swiping up or down, or gesture operations. In response to the first trigger action, the current playback time of the video is determined, and the nearest highlight scene to the current playback time is selected from the highlight scene information as the target highlight scene.

[0216] like Figure 15 As shown, in response to the first triggering operation, such as a swipe operation, the time jump is initiated from the time point of story 1 to the time point of story 2.

[0217] The first trigger operation includes a forward jump operation and / or a backward jump operation. If a forward jump operation is received, the system searches for the previous time point in the featured story information as the target featured story segment. For example, based on a received downward swipe operation, the system searches for the previous time point in the featured story information as the target featured story segment, and playback begins from that target featured story segment.

[0218] If a jump-back command is received, the system searches for the next available highlight segment in the highlight segment information, and selects that segment as the target highlight segment. For example, based on a received swipe-up command, the system searches for the next available highlight segment in the highlight segment information, selects that segment as the target highlight segment, and begins playback from that target highlight segment.

[0219] Step 1406: Jump the target video to the time point corresponding to the exciting plot segment and play it.

[0220] Obtain the video data corresponding to the target exciting plot segment, parse and render the video data, and jump to the video data rendered by the target exciting plot segment for playback.

[0221] In this embodiment of the invention, the target video can jump between various exciting plot segments. Responding to a second trigger operation, playback jumps between exciting plot segments. The second trigger operation is the operation that triggers the jump between exciting plot segments; the first trigger operation and the second trigger operation can be the same operation or different operations, and this embodiment of the invention does not impose any restrictions. Upon receiving the second trigger operation, playback jumps between exciting plot segments in response to the second trigger operation. If the second trigger operation is a forward jump operation (or playback of the previous plot segment), then the previous exciting plot segment of the current exciting plot segment is determined, and playback jumps to begin from that previous exciting plot segment. If the second trigger operation is a backward jump operation (or playback of the next plot segment), then the next exciting plot segment of the current exciting plot segment is determined, and playback jumps to begin from that next exciting plot segment. This enables video data to jump between various plot segments, improving the user's video viewing experience.

[0222] This approach preserves the complete narrative of long videos while allowing users to enjoy a viewing pace similar to short videos, effectively improving user experience and platform retention.

[0223] Reference Figure 17 The diagram illustrates a flowchart of another embodiment of the video playback method based on key points according to the present invention.

[0224] Step 1702: Play the target video on the playback page. The target video is associated with exciting plot information.

[0225] Step 1704: In response to the first trigger operation, determine the most recent exciting plot segment based on the exciting plot information.

[0226] Step 1706: Jump the target video to the time point corresponding to the exciting plot segment and play it.

[0227] Step 1708: Display the title of the exciting plot segment on the playback page.

[0228] If the server receives a playback request from the client, it can send the target video data to the client. The client then parses and renders the video data and plays the target video on the playback page. The client receives a first trigger operation, such as an up / down swipe or a gesture. In response to this first trigger operation, if a "jump forward" operation is received, the server searches for the first relevant scene segment before the current playback time in the scene segment information and uses it as the target scene segment. If a "jump backward" operation is received, the server searches for the first relevant scene segment after the current playback time in the scene segment information and uses it as the target scene segment.

[0229] The system retrieves the video data corresponding to the target exciting plot segment, parses and renders the video data, redirects the target video to start playing from the video data rendered from that target exciting plot segment, and displays the title on the playback page. For example... Figure 15 As shown, in response to the first trigger operation, the playback jumps from the time point of plot 1 to the time point of plot 2 and displays the title of plot 2 on the playback page: A Discovers the Spirit Stone.

[0230] Based on the above embodiments, this invention provides a video playback system, such as... Figure 18 As shown, the system includes: a server 1802 and a client 1804, wherein:

[0231] Server 1802 obtains playback information corresponding to the target video, the playback information including: video segment information divided according to playback speed, the playback speed including double speed playback and normal speed playback; based on the playback speed, determines the switching video segment corresponding to the target switching behavior, and based on the switching video segment corresponding to the target switching behavior, determines multiple exciting video segments; adjusts the multiple exciting video segments based on audio and video content to determine multiple exciting plot information, the exciting plot information including the time points corresponding to the exciting video segments; and establishes association between the multiple exciting plot information and the target video.

[0232] Client 1804 plays the target video on the playback page and, in response to the first trigger operation, determines the most recent exciting plot segment based on the exciting plot information; then jumps the target video to the time point corresponding to the exciting plot segment and plays it.

[0233] The video key point extraction and playback method of this invention is of significant value to long-form video services on video platforms. It innovatively solves the problem of users encountering uninteresting content while watching long videos. When users feel that the current plot is not engaging or relatable, they no longer need to exit the video or aimlessly fast-forward. Instead, they can easily jump to the next exciting plot segment by simply swiping in the left or right quarter of the screen. This "second chance" mechanism effectively reduces user abandonment rates because users are more inclined to try jumping to other exciting segments to satisfy their viewing needs than to exit the video directly.

[0234] This invention also provides an electronic device, such as... Figure 19 As shown, it includes a processor 191, a communication interface 192, a memory 193, and a communication bus 194, wherein the processor 191, the communication interface 192, and the memory 193 communicate with each other through the communication bus 194.

[0235] Memory 193 is used to store computer programs;

[0236] When processor 191 executes a program stored in memory 193, it performs the following steps:

[0237] Obtain the playback speed of each video segment in the target video;

[0238] Based on the playback speed, determine the switching video segment corresponding to the target switching behavior; based on the switching video segment corresponding to the target switching behavior, determine multiple exciting video segments and exciting plot information, wherein the exciting plot information includes the time points corresponding to the exciting video segments.

[0239] The multiple pieces of exciting plot information are associated with the target video, so that when the target video is played on the playback page, in response to the first trigger operation, the exciting plot segment closest to the current time point is determined based on the exciting plot information, and the playback is initiated from the target video to the time point corresponding to the exciting plot segment.

[0240] In another alternative embodiment, when the processor 191 executes the program stored in the memory 193, it performs the following steps:

[0241] Playing a target video on the playback page, the target video is associated with exciting plot information, the exciting plot information includes: the time point corresponding to the exciting plot segment, the exciting plot segment is determined based on the audio and video content adjustment processing of the exciting video segment, the exciting video segment is determined based on the switching video segment corresponding to the target switching behavior, and the target switching behavior is determined based on the playback speed of each video segment in the target video.

[0242] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0243] The communication interface is used for communication between the aforementioned terminal and other devices.

[0244] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0245] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0246] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform the video key point extraction method and the key point-based video playback method described in any of the above embodiments.

[0247] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the video key point extraction method and the key point-based video playback method described in any of the above embodiments.

[0248] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0249] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0250] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0251] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A method for extracting key points from a video, characterized in that, The method includes: Obtain the playback speed of each video segment in the target video; Based on the playback speed, determine the switching video segment corresponding to the target switching behavior; based on the switching video segment corresponding to the target switching behavior, determine multiple exciting video segments and exciting plot information, wherein the exciting plot information includes the time points corresponding to the exciting video segments. The multiple pieces of exciting plot information are associated with the target video, so that when the target video is played on the playback page, in response to the first trigger operation, the exciting plot segment closest to the current time point is determined based on the exciting plot information, and the playback is initiated from the target video to the time point corresponding to the exciting plot segment.

2. The method according to claim 1, characterized in that, The playback speed includes double-speed playback and normal speed playback, and the double-speed playback includes fast playback and slow playback; determining the switching video segment corresponding to the target switching behavior based on the playback speed includes: Determine the time point corresponding to the target switching behavior, and obtain a switching video segment of a predetermined duration starting from the time point. The target switching behavior includes: a first behavior of switching from fast playback to normal speed playback or slow playback, and / or a second behavior of switching from normal speed playback to slow playback.

3. The method according to claim 2, characterized in that, The step of determining multiple exciting video clips and plot information based on the target switching behavior and corresponding video segments includes: Based on the video clips switched by each user, the number of overlaps at time points is counted, and a playback popularity curve of the target video is generated based on the number of overlaps. Based on the aforementioned popularity curve, several highlight video clips were identified. Extract the start and end times of exciting video clips as key plot information.

4. The method according to claim 3, characterized in that, The determination of multiple highlight video clips based on the playback popularity curve includes: Determine the average value of the playback popularity corresponding to the playback popularity curve; Based on the average value, a candidate time series corresponding to the popularity peak of the playback popularity curve is determined, and the candidate time series includes the timestamps of candidate time points where the popularity peak is higher than the average value; The candidate time series are filtered according to a predetermined time window to determine multiple target time points; Obtain multiple exciting video clips corresponding to the multiple target time points.

5. The method according to claim 1, characterized in that, Also includes: The key video clips were adjusted based on the audio and video content to determine the adjusted key plot information.

6. The method according to claim 5, characterized in that, The process of adjusting the key video clips based on audio and video content to determine the adjusted key plot information includes: The video clips are subjected to image detection to determine the key plot information of the key plot segments starting from specific scenes; and / or The exciting video clips are subjected to dialogue coherence detection to determine the exciting plot information of the exciting plot clips starting from the dialogue entry point. The dialogue entry point includes: the dialogue end point and / or the dialogue start point.

7. The method according to claim 5, characterized in that, The process of adjusting the key video clips based on audio and video content to determine the adjusted key plot information includes: Based on a pre-built language model, the exciting video clips are analyzed from multiple dimensions to understand the plot, and the adjusted plot information is output.

8. The method according to claim 7, characterized in that, The process of performing multi-dimensional plot understanding on the exciting video clips based on a pre-built language model and outputting the adjusted plot information includes: Obtain the video description information corresponding to the exciting video clip, the video description information including the audio and video content corresponding to the exciting video clip; Obtain a prompt word template, which includes variable parameters and prompt items. The prompt items include at least one of the following: analysis content, evaluation priority, and output requirements. The video description information and the prompt word template are combined to generate input information; The input information is fed into a pre-set language model for multi-dimensional plot understanding, and the adjusted plot information is output, which includes the time information and title of the plot segment.

9. The method according to claim 4, characterized in that, Also includes: For two adjacent target time points whose time interval exceeds a first time threshold, determine the video description information of the video segment to be detected between the two adjacent target time points, wherein the video description information includes the audio and video content of the video segment to be detected; Based on the audio-visual content, the video segment to be detected is analyzed from multiple dimensions to determine key plot information.

10. The method according to claim 4, characterized in that, Also includes: For two adjacent target time points whose time interval does not exceed the second time threshold, the video segments between the two adjacent target time points are merged to generate the video segment to be detected. Based on the audio-visual content, the video segment to be detected is analyzed from multiple dimensions to determine key plot information.

11. The method according to claim 1, characterized in that, The process of associating various key plot details with the target video includes: Jump points are set in the target video according to the time points of the exciting plot segments, so as to jump between the jump points in response to the first trigger operation.

12. A video playback method based on key points, characterized in that, The method includes: Play a target video on the playback page. The target video is associated with exciting plot information, which includes: the time point corresponding to the exciting plot segment. The exciting plot segment is determined based on the switching video segment corresponding to the target switching behavior. The target switching behavior is determined based on the playback speed of each video segment in the target video. In response to the first trigger operation, the most recent exciting plot segment is determined based on the exciting plot information; Jump to the target video and play it at the corresponding time point of the exciting plot segment.

13. The method according to claim 12, characterized in that, The first triggering operation includes a forward jump operation and / or a backward jump operation. The response to the first triggering operation, determining the most recent exciting plot segment based on the exciting plot information, includes: If a forward jump operation is received, the exciting plot segment corresponding to the previous time point of the current playback time point is searched in the exciting plot information. If a jump-back command is received, the program will search for the next time point in the featured plot information to find the featured plot segment corresponding to the current playback time point.

14. The method according to claim 12, characterized in that, Also includes: In response to the second triggered action, the playback jumps between exciting plot clips.

15. The method according to claim 12, characterized in that, Also includes: When the terminal device is in landscape mode, at least one jump operation area is set in the playback page; The first triggering operation includes a sliding operation received in the jump operation area.

16. The method according to claim 12, characterized in that, Also includes: When the playback jumps to the specified time point, the title of the exciting plot segment is displayed on the playback page.

17. A video playback system, characterized in that, The system includes: a server and a client; The server obtains the playback speed of each video segment in the target video; determines the switching video segment corresponding to the target switching behavior based on the playback speed; determines multiple exciting video segments and exciting plot information based on the switching video segments corresponding to the target switching behavior, wherein the exciting plot information includes the time points corresponding to the exciting video segments; and establishes associations between the multiple exciting plot information and the target video. The client plays the target video on the playback page and, in response to the first trigger operation, determines the most recent exciting plot segment based on the exciting plot information; then jumps the target video to the time point corresponding to the exciting plot segment and plays it.

18. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-16.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-16.

20. A computer program product comprising a computer program / computer-executable instructions, wherein, When the computer program / computer executable instructions are executed by a processor in an electronic device, they implement the method described in any one of claims 1-16.

Citation Information

Cited By

  • IPTV program time period analysis method and system

    CN121357373A

  • An iptv program time slot analysis method and system

    CN121357373B