Video skip point generation method and device, equipment and medium

By acquiring dialogue information, video descriptions, and user behavior data, video segments are segmented according to program type, and interest points are analyzed using a large model. This solves the problem that existing technologies cannot adapt to different program types, achieves exciting and accurate segmentation of skip points, and improves the user experience.

CN122496683APending Publication Date: 2026-07-31BEIJING IQIYI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING IQIYI TECH CO LTD
Filing Date
2026-05-25
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technology cannot adapt to variety show videos of different program types, resulting in video segments that do not match users' attention span and interests, and the generated skip points are not exciting enough.

Method used

By acquiring dialogue information, video descriptions, and user behavior data, video segments are segmented according to program type. A large model is used to analyze points of interest and generate jump points to adapt to different program types, ensuring that the segmented video segments match user interests and attention spans.

Benefits of technology

It improves the accuracy of jump view point segmentation, ensuring that the generated video clips are engaging and match user interests, thus enhancing the user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496683A_ABST
    Figure CN122496683A_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, device, and medium for generating video skip points. The method includes: acquiring a video to be segmented and its video data; the video data includes at least: dialogue information, video description, and user behavior data; segmenting the video to be segmented into video segments of different program types based on the dialogue information and video description; determining the user's points of interest in the video segments based on the user behavior data; generating skip points for the video to be segmented based on the dialogue information, video description, and points of interest; the skip points represent the client responding to an adjustment command, jumping the playback progress of the video to be segmented to the start time of the target video segment. Segmenting video segments by program type solves the problem that video segments obtained by scene division cannot adapt to different program types; determining the user's points of interest based on user behavior data and segmenting the video segments according to those points of interest ensures the quality of the content played at the generated skip points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet application technology, and in particular to a method, apparatus, device, and medium for generating video jump view points. Background Technology

[0002] With the rise of short videos, more and more users are accustomed to watching shorter videos, and their attention span is changing accordingly, causing them to focus their attention on certain segments. Variety shows, on the other hand, are longer, so to better adapt to this shift in user attention, related technologies typically segment the video according to scenes, allowing viewers to watch video segments based on those scenes.

[0003] However, the skip points obtained by segmenting variety show videos according to the scene cannot be adapted to variety show videos of different program types, and cannot ensure that the segmented video segments are exciting segments that match the user's attention span. There are problems such as incompatibility with different program types and inaccurate segmentation of skip points. Summary of the Invention

[0004] The purpose of this invention is to provide a method, apparatus, device, and medium for generating video skip points, adapting to different video program types and improving the accuracy of skip point segmentation for different video program types. The specific technical solution is as follows: In a first aspect of this invention, a method for generating video skip points is provided, the method comprising: Obtain the video to be segmented and its video data; the video data includes at least: dialogue information, video description, and user behavior data; Based on the dialogue information and the video description, the video to be segmented is divided into video segments of different program types; Determine the user's points of interest in the video clip based on the user behavior data; The jump points of the video to be segmented are generated based on the dialogue information, the video description, and the points of interest; the jump points indicate that the client responds to the adjustment command and jumps the playback progress of the video to be segmented to the start time point of the target video segment.

[0005] Optionally, the step of segmenting the video to be segmented into video segments of different program types based on the dialogue information and the video description includes: The program types contained in the video to be segmented are determined based on the dialogue information and the video description. Match the corresponding video segmentation strategy according to the program type; The video to be segmented is divided into multiple video segments according to the video segmentation strategy and the program type.

[0006] Optionally, the program type includes at least: performance programs and talk shows; the video segmentation strategy includes: a first video segmentation strategy and a second video segmentation strategy; The step of dividing the video to be divided into multiple video segments according to the video segmentation strategy and the program type includes: If the video to be segmented contains performance programs, a first video segmentation strategy is used to extract the first starting time point of each performance program in the video to be segmented; and the video to be segmented is segmented into video segments containing performance programs based on the first starting time point. If the video to be segmented contains talk shows, the second starting time point of each talk show in the video to be segmented is extracted according to the second video segmentation strategy. The second starting time point, the dialogue information and the video synopsis are then input into a pre-trained large model to determine the structured segments of the talk shows. Based on the second starting time point and the structured segments, the video to be segmented is segmented into video segments containing the talk shows.

[0007] Optionally, the user behavior data includes at least: a playback curve and a video speed adjustment point; the video speed adjustment point represents the maximum value point at which the user adjusts the video to be segmented from the first speed to the second speed, wherein the first speed is greater than the second speed; Determining the user's interest in the video clip based on the user behavior data includes: The maximum value point of the playback curve is determined based on the playback curve. The point of interest is obtained by fusing the maximum value point during playback with the video speed adjustment point.

[0008] Optionally, generating the skip points for the video to be segmented based on the dialogue information, the video description, and the points of interest includes: The video segment is further segmented based on the points of interest and a preset attention duration threshold to obtain multiple target video segments. The jump view points of the video to be segmented are determined based on the dialogue information, the video description, and the start time of the target video segment.

[0009] Optionally, the step of further segmenting the video segment based on the points of interest and a preset attention duration threshold to obtain multiple target video segments includes: If the duration of the video segment exceeds the preset attention duration threshold, the video segment is further segmented based on the interest points to obtain a target video segment; the target video segment contains the interest points. If the duration of the video segment is less than or equal to the preset attention duration threshold, the video segment is used as the target video segment.

[0010] Optionally, the method further includes: Generate video script for the target video segment based on the dialogue information and video description contained in the target video segment; Obtain the user's adjustment command for the preset switching area on the playback page of the video to be segmented; Upon receiving the adjustment instruction, the playback progress of the video to be segmented is redirected to the nearest skip point, the target video segment is played, and the video text is displayed.

[0011] In a second aspect of the present invention, a video skipping point generation device is also provided, the device comprising: The video to be segmented acquisition module is used to acquire the video to be segmented and the video data of the video to be segmented; the video data includes at least: dialogue information, video description and user behavior data; The video segmentation module is used to segment the video to be segmented into video segments of different program types based on the dialogue information and the video description. An interest point generation module is used to determine the user's interest points in the video clip based on the user behavior data; The jump point generation module is used to generate jump points for the video to be segmented based on the dialogue information, the video description, and the points of interest; the jump point indicates that the client responds to the adjustment command and jumps the playback progress of the video to be segmented to the start time point of the target video segment.

[0012] In another aspect of the present invention, a computer-readable storage medium is also provided, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform any of the video jump point generation methods described above.

[0013] In another aspect of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the video jump point generation methods described above.

[0014] The video skip point generation method provided in this invention involves acquiring the video to be segmented and its video data. The video data includes at least dialogue information, video description, and user behavior data. The video to be segmented is divided into video segments of different program types based on the dialogue information and video description. User interest in the video segments is determined based on the user behavior data. Skip points are generated for the video to be segmented based on the dialogue information, video description, and points of interest. Each skip point represents a client's response to an adjustment command, causing the playback progress of the video to be segmented to jump to the start time of the target video segment. Segmenting video segments by program type solves the problem that video segments obtained based on scene division cannot adapt to different program types. Determining user interest based on user behavior data and segmenting video segments according to that interest ensures the quality of the content played at the generated skip points. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0016] Figure 1 This is a flowchart of the steps of a video skip point generation method provided in an embodiment of the present invention; Figure 2 This is a screenshot of the playback page in a video jump point generation method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a video skip point generation method device provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of a computer-readable storage medium provided in an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0018] With the rise of short videos, more and more users are accustomed to watching shorter videos, and their attention span is changing accordingly, causing them to focus their attention on certain segments. Variety shows, on the other hand, are longer. To better adapt to this shift in user attention, related technologies typically segment the video according to scenes or summary information, generating corresponding skip points, which then allow users to jump to and watch different video segments.

[0019] For example, the prior art patent number CN119271393A discloses a method for generating video highlights. This technology determines the skip points based on the single episode summary information of the target video, as well as the text information and scene information appearing in the target video. Then, the user adjusts the viewing highlights by dragging the progress bar to the skip point. However, the prior art cannot adapt to videos of different program types and cannot ensure that the segmented video segments are in line with the user's interests. As a result, the content played after the generated skip point does not match the user's interests and is not exciting enough.

[0020] Reference Figure 1 The diagram illustrates a flowchart of a video skip point generation method provided in an embodiment of the present invention, which may specifically include the following: Step S101: Obtain the video to be segmented and its video data; the video data includes at least: dialogue information, video description, and user behavior data; In this embodiment, the video to be segmented can be different types of videos on a video platform. The video platform will collect video data for the video, which includes at least dialogue information, video description and user behavior data. The user behavior data includes data on different users' fast forwarding, rewinding and playback speed adjustment of the video.

[0021] Step S102: Divide the video to be segmented into video segments of different program types according to the dialogue information and the video description; In practice, the videos to be segmented can be divided into different program types. Taking variety shows as an example, variety shows often contain different types of content, including singing content as well as talk shows, interviews and other language-based content. Due to the differences in performance styles, the video segmentation strategies for singing content and language-based content are different.

[0022] This embodiment determines the specific program type based on the dialogue information and video description, and then uses the corresponding segmentation strategy to segment it into video segments. Step S103: Determine the user's points of interest in the video clip based on the user behavior data; In this embodiment, the user behavior data is a statistical analysis of user playback behaviors at different times in the video to be segmented, which may include user actions such as fast forwarding, rewinding, and adjusting playback speed.

[0023] Specifically, if the video to be segmented is dragged to a later time point, that time point will be counted once in the replay points; if at any time point of the video to be segmented, the user adjusts the playback speed from high speed to low speed or normal speed, it means that there is content that the user likes to watch at that time point. By analyzing this user behavior data, we can determine the points of interest that match the user's viewing habits.

[0024] Step S104: Generate the jump point of the video to be segmented based on the dialogue information, the video description and the points of interest; the jump point indicates that the client responds to the adjustment command and jumps the playback progress of the video to be segmented to the start time point of the target video segment.

[0025] In this embodiment, the video segments are video content of different program types. By analyzing the dialogue information, video description and points of interest, at least one target video segment containing points of interest and complete content can be generated. Jump points are generated according to the start time of the target video segment in the video to be segmented, so that the client can respond to the user's adjustment command and jump the playback progress of the video to be segmented to the start time of the target video segment.

[0026] Specifically, in this embodiment, a large model can be pre-trained, taking dialogue information, video description, and points of interest as input. This allows the large model to segment a complete video containing points of interest based on the input information.

[0027] In this embodiment of the invention, the method for generating video skip points includes: acquiring the video to be segmented and its video data; the video data includes at least: dialogue information, video description, and user behavior data; segmenting the video to be segmented into video segments of different program types based on the dialogue information and video description; determining the user's points of interest in the video segments based on the user behavior data; generating skip points for the video to be segmented based on the dialogue information, video description, and points of interest; the skip points represent the client's response to an adjustment command, jumping the playback progress of the video to be segmented to the start time of the target video segment. Segmenting video segments by program type solves the problem that video segments obtained by scene division cannot adapt to different program types; determining the user's points of interest based on user behavior data and segmenting the video segments according to those points of interest ensures the quality of the content played at the generated skip points.

[0028] In one embodiment of the present invention, the step of segmenting the video to be segmented into video segments of different program types according to the dialogue information and the video description includes: The program types contained in the video to be segmented are determined based on the dialogue information and the video description. Match the corresponding video segmentation strategy according to the program type; The video to be segmented is divided into multiple video segments according to the video segmentation strategy and the program type.

[0029] In this embodiment, the video to be segmented contains video content of different program types. The program type contained in the video to be segmented can be determined based on the dialogue information and video synopsis (e.g., plot summary). Then, the corresponding video segmentation strategy is adopted to segment the video to be segmented into multiple video segments based on the program type contained in the video. Each video segment corresponds to the video content contained in a program type.

[0030] In order to improve the efficiency of segmenting skip points and the viewing experience of the segmented target videos, the videos to be segmented can be obtained by intelligent sampling and compression of the original videos of variety shows.

[0031] Specifically, motion intensity analysis can be performed on the video frames of the original video to determine the changes between different video frames. Then, the sampling frequency at different time points can be adjusted according to the changes between different video frames. After that, key frames of the original video are extracted according to the sampling frequency. By collecting these key frames, a compressed video is generated to compress the size of the original video. Finally, the resolution of the corresponding video content is adjusted according to the video content of the compressed video to obtain the video to be segmented.

[0032] In practical implementation, the motion intensity of the original video can be analyzed based on optical flow. This involves reading video frames and converting them to grayscale images. By calculating the Farneback optical flow of the previous and current frames, the motion field of the entire image is calculated for global intensity analysis. Then, optical flow vector decomposition is performed, transforming the optical flow (dx, dy) into polar coordinates (magnitude, angle). The magnitude is used to calculate the average motion intensity. When the average motion intensity exceeds a preset threshold, the motion intensity is considered high. For time points with high motion intensity, the sampling rate is increased, meaning frames are captured more densely to preserve details. For time points with low motion intensity, the sampling rate is decreased to reduce redundant data. Through this dynamic adjustment, the frame sequence is intelligently optimized. While ensuring the complete extraction of key motion information, the temporal dimension is effectively compressed, resulting in a compressed video with fewer frames but more coherent content.

[0033] Secondly, this invention determines different resolutions based on the program type of variety show videos. For example, for live-action programs, a high resolution is used, while for talk show programs, since this type of program focuses more on language expression and has lower requirements for picture quality, the resolution can be reduced.

[0034] Optionally, keyframes of the original video are extracted based on the sampling frequency to generate a compressed video, including: Based on the sampling frequency, select video frames from the original video and determine the histogram of the video frames; If the Bach distance between the histograms of adjacent video frames is greater than or equal to a distance threshold, the video frame is designated as the keyframe. Extract each keyframe from the original video to generate the compressed video.

[0035] In this embodiment, after determining the sampling rate at different time points based on the motion intensity, video frames of the original video are selected according to the obtained sampling rate, the histogram of each video frame is calculated, and the Bhattacharyya distance is used to compare the histograms of adjacent frames. If the Bhattacharyya distance between the histograms of adjacent video frames is greater than or equal to the distance threshold, the video frame is used as a key frame. Then, by extracting each key frame of the original video, the final compressed video is generated.

[0036] By adaptively and dynamically adjusting the frame sampling rate based on motion intensity and selecting video frames for keyframe determination, the extraction accuracy and compression efficiency of exciting clips can be significantly improved, ensuring the complete preservation of key content while reducing storage and transmission costs.

[0037] In one embodiment of the present invention, the program type includes at least: performance programs and talk shows; the video segmentation strategy includes: a first video segmentation strategy and a second video segmentation strategy; The step of dividing the video to be divided into multiple video segments according to the video segmentation strategy and the program type includes: If the video to be segmented contains performance programs, a first video segmentation strategy is used to extract the first starting time point of each performance program in the video to be segmented; and the video to be segmented is segmented into video segments containing performance programs based on the first starting time point. If the video to be segmented contains talk shows, the second starting time point of each talk show in the video to be segmented is extracted according to the second video segmentation strategy. The second starting time point, the dialogue information and the video synopsis are then input into a pre-trained large model to determine the structured segments of the talk shows. Based on the second starting time point and the structured segments, the video to be segmented is segmented into video segments containing the talk shows.

[0038] In this embodiment, the program type can be divided into performance programs and talk shows; corresponding video segmentation strategies are adopted for different program types.

[0039] In practice, performance programs include at least song performances or dance performances. If the singing programs are segmented, it will result in the inability to watch the complete performance content, and there will be gaps between the songs or dance videos, which will damage the user's viewing and listening experience.

[0040] Therefore, when the video to be segmented contains at least one song performance or dance performance, the first video segmentation strategy is used to extract the first starting time point of each song performance or dance performance in the video to be segmented. The first starting time point represents the starting time point of the video content of the performance in the video to be segmented. Then, according to the first starting time point, each song performance or dance performance in the video to be segmented is segmented into a corresponding video segment. That is, each song or each dance performance is set as a separate video segment, and each video segment contains the video content of a complete song performance or dance performance.

[0041] In this embodiment, talk shows include at least interview shows (e.g., talk shows and interview programs) and hosting shows. The video content of interview shows and hosting shows is usually long and contains a lot of dialogue content with low engagement (e.g., character introductions, rule introductions, and advertising recommendations in interview shows and hosting shows). Users are usually not interested in this low engagement dialogue content and find it difficult to watch this type of program in its entirety. Therefore, it is necessary to further segment interview shows and hosting shows.

[0042] Specifically, when the video to be segmented contains talk shows, a second video segmentation strategy is used to extract the second starting time point of each talk show in the video to be segmented. The second starting time point is the starting time point of the video content of the talk show in the video to be segmented. Then, the second starting time point, the dialogue information and the video synopsis are input into a pre-trained large model, and the large model determines the structured segments of the talk shows. Finally, based on the second starting time point and the structured segments, the video to be segmented is segmented into multiple video segments containing structured segments.

[0043] Since talk shows are primarily conversational programs, users are more interested in video clips with dense dialogue and trending topics. Based on this characteristic, video clips with dense dialogue and trending topics can be used as structured segments for talk shows. Specifically, video segments with dense dialogue information can be selected from the video content included at the second starting time point based on the dialogue information. Hot topics involved in talk shows can be determined based on the video descriptions. Then, content containing hot topics can be selected from the video segments with dense dialogue information and set as structured segments.

[0044] In practical applications, the historical structured segments of talk shows and the corresponding dialogue information and video summaries can be used as training data. The model can be trained using this training data, so that the trained large model can learn the mapping relationship between structured segments and dialogue information and video summaries. Then, the large model can analyze the structured segments of talk shows in the video to be segmented in real time.

[0045] Optionally, since talk shows focus more on language expression and have lower requirements for visual quality, the resolution of the talk show video can be reduced. The resolution of the current video content of the talk show can be adjusted to a preset resolution, which is lower than the current resolution of the talk show. The adjusted video content, dialogue information, and video synopsis of the talk show are then input into a pre-trained large model, which can further improve the efficiency of the large model in analyzing exciting segments.

[0046] This embodiment implements a differentiated processing strategy based on program type. For performance programs, no segmentation is performed; the complete segments of each performance program are directly retained to avoid content fragmentation and ensure a good viewing experience. For talk shows such as interviews and hosting programs, segmentation is performed using a large model to accurately identify structured segments that users are more interested in and effectively eliminate redundant and low-value segments. This differentiated processing strategy can significantly improve the coherence of video segment extraction and ensure the quality of video content.

[0047] In one embodiment of the present invention, the user behavior data includes at least: a playback curve and a video speed adjustment point; the video speed adjustment point represents the maximum value point at which the user adjusts the video to be segmented from a first speed to a second speed, wherein the first speed is greater than the second speed; Determining the user's interest in the video clip based on the user behavior data includes: The maximum value point of the playback curve is determined based on the playback curve. The point of interest is obtained by fusing the maximum value point during playback with the video speed adjustment point.

[0048] In practical applications, user behavior data refers to statistics that reflect user playback behavior at different times in the video to be segmented, including at least: a playback curve and video speed adjustment points. The playback curve represents the playback statistics at different time points. For example, if a user drags the video to be segmented forward to a certain time point, this behavior can be counted as a playback at that time point. The video speed adjustment point represents the maximum value at which the user adjusts the playback speed of the video to be segmented from the first speed to the second speed at any time point, where the first speed is greater than the second speed. For example, the first speed can be understood as high speed, and the second speed can be understood as low speed or normal speed.

[0049] Specifically, the maximum playback point in the playback curve can be determined based on the playback curve; then, the maximum playback point can be merged with the video speed adjustment point to obtain the user's points of interest for the video to be segmented, and these points of interest can be used as the basis for the final segmentation of the second video segment.

[0050] In the specific implementation, if the time interval between the maximum point of playback and the video speed adjustment point is less than or equal to the preset interval threshold, the maximum point of playback and the video speed adjustment point can be merged into a fusion point. Before the fusion point, an interest point is selected to ensure that the target video segment obtained from the interest point can contain complete content and will not be segmented from the middle of the dialogue. If the time interval between the maximum value point and the video speed adjustment point is greater than the preset interval threshold, set an interest point before the maximum value point and another interest point before the video speed adjustment point to ensure that no exciting content is missed.

[0051] In practical applications, an interval threshold can be pre-set based on historical data of the user's replay of the maximum value point and the video speed adjustment point (for example, it can be set to 30 seconds). If the replay of the maximum value point is at 10 seconds in the video clip and the video speed adjustment point is at 15 seconds in the video clip, then the time interval between these two time points is 5 seconds. This time interval is less than the interval threshold, so it can be determined that the user is more interested in the video content after 10 seconds in the video clip. Therefore, the earlier time points of the replay of the maximum value point and the video speed adjustment point in the video clip can be set as the fusion point. Before the fusion point, an interest point that can contain complete content and will not be cut from the middle of the dialogue can be selected as the basis for the final segmentation. If the maximum point in the video clip is at 10 seconds, and the video speed adjustment point is at 42 seconds, then the time interval between these two points is 32 seconds. This time interval is greater than the interval threshold. In this case, the user may be interested in the video content after 10 seconds and the video content after 42 seconds, respectively. Therefore, it is necessary to select a point of interest before the maximum point that contains complete content and is not interrupted by the dialogue, and a point of interest before the video speed adjustment point that contains complete content and is not interrupted by the dialogue, to ensure that no exciting content is missed.

[0052] This embodiment analyzes the statistical results of user playback behavior at different times, dynamically integrates the maximum replay value point with the speed adjustment point, achieves accurate positioning of the point of interest, effectively avoids cutting from the middle of the dialogue, ensures the integrity and coherence of the target video segment, significantly improves the semantic coherence of exciting segments and the user viewing experience, and provides data support for the subsequent generation of skip points.

[0053] In one embodiment of the present invention, generating the skip points of the video to be segmented based on the dialogue information, the video description, and the points of interest includes: The video segment is further segmented based on the points of interest and a preset attention duration threshold to obtain multiple target video segments. The jump view points of the video to be segmented are determined based on the dialogue information, the video description, and the start time of the target video segment.

[0054] In this embodiment, the preset attention duration threshold can be set as the attention duration of users of different program types obtained in advance. When the duration of a video segment is longer than the preset attention duration threshold, the duration of the video segment cannot ensure that the user can watch it completely. Therefore, it is necessary to perform secondary segmentation of the video segment according to the points of interest and the preset attention duration threshold to obtain the target video segment. Next, the target video clip is combined with the dialogue information and video description to determine the jump points that can contain the target video clip and whose video content is complete. This ensures that the target video clip played from the jump point can meet both the video's level of excitement and the user's attention span.

[0055] In one embodiment of the present invention, the step of performing secondary segmentation of the video segment based on the interest point and a preset attention duration threshold to obtain multiple target video segments includes: If the duration of the video segment exceeds the preset attention duration threshold, the video segment is further segmented based on the interest points to obtain a target video segment; the target video segment contains the interest points. If the duration of the video segment is less than or equal to the preset attention duration threshold, the video segment is used as the target video segment.

[0056] In practice, as users' attention span changes, video clips may sometimes exceed their attention span. If a video clip exceeds this length, users may lose focus and miss or skip the portion that exceeds their attention span, leading them to perceive the variety show as less engaging and compromising their overall experience by not receiving the complete and exciting content.

[0057] To ensure that the duration of the target video segment corresponding to the jump view point does not exceed the user's attention span, this embodiment uses the point of interest and the preset attention span threshold as the segmentation basis to perform secondary segmentation of the video segment, thereby obtaining the target video segment and the corresponding jump view point.

[0058] In practical implementation, since the attention span of users may vary for different program types, this invention can set a preset attention span threshold for the program type by statistically analyzing the completion rate of videos of different program types at different durations. This adapts to the changes in user attention span and ensures that the second video segment obtained by segmentation will attract users.

[0059] If the duration of a video clip exceeds the preset attention duration threshold, the duration of the video clip does not meet the user's attention duration limit. This may result in the user being unable to concentrate on watching the complete video clip, missing or skipping the part of the video clip that exceeds the attention duration. Therefore, it is necessary to perform a secondary segmentation of the video clip based on the point of interest, so that the duration of the segmented target video clip is less than or equal to the preset attention duration threshold, and the starting time of the target video clip includes the point of interest.

[0060] In practical applications, when the duration of a video segment exceeds a preset attention duration threshold, the video segment and points of interest can be input into a pre-trained large model to identify each video sub-segment that meets the preset attention duration threshold. Then, it is determined whether the video sub-segment contains points of interest, and the video sub-segments are filtered. Finally, the target video segment with complete video content is selected from the filtered video sub-segments.

[0061] If the duration of a video clip is less than or equal to a preset attention duration threshold, the duration of the video clip meets the user's attention duration limit. Therefore, the video clip can be used as the target video clip, and jump points can be generated based on the start time of the target video clip in the video to be segmented.

[0062] This embodiment determines a preset attention duration threshold based on statistics of user attention duration for different program types. It correlates video segmentation with user attention duration, performing targeted segmentation based on the relationship between video segment length and user attention duration. When a video segment's length exceeds the preset threshold, it is precisely segmented based on points of interest, generating target segments and skip points that match the user's attention cycle. When a video segment's length is less than or equal to the threshold, it is directly retained as a complete target segment, and skip points are generated. This method effectively adapts to changes in user attention duration, ensuring that the target segment length accurately matches user viewing habits, avoiding viewing interruptions due to excessively long content, and significantly improving content continuity and user satisfaction. Simultaneously, the duration threshold constraint optimizes segment length and eliminates redundant content.

[0063] In one embodiment of the present invention, the method further includes: Generate video script for the target video segment based on the dialogue information and video description contained in the target video segment; Obtain the user's adjustment command for the preset switching area on the playback page of the video to be segmented; Upon receiving the adjustment instruction, the playback progress of the video to be segmented is redirected to the nearest skip point, the target video segment is played, and the video text is displayed.

[0064] In this embodiment, in addition to segmenting the target video segments based on points of interest and a preset attention duration threshold set based on user attention duration, a video script for the target video segment is generated based on the dialogue information and video description contained in the target video segment. When a user's adjustment instruction is obtained (e.g., a swipe operation performed in a preset switching area on the playback page of the video to be segmented, recognizing the user's intention to adjust the video playback progress), the current playback progress of the video to be segmented is jumped to the nearest jump point, the target video segment is played, and the video script of the target video segment is displayed so that the user can quickly understand the content summary of different target video segments.

[0065] Reference Figure 2 This image shows the effect of the playback page in a video jump point generation method provided in an embodiment of the present invention. Figure 2As can be seen, this invention divides the playback page of the video to be segmented into three functional areas. A volume adjustment area and a brightness adjustment area are set in the middle of the playback page, used to adjust the volume and screen brightness of the video to be segmented, respectively. On the left and right sides of the playback page, operation areas for switching target video segments are set. By obtaining the user's up and down adjustment commands in these areas, the target video segments can be switched. Unlike traditional methods that rely on users dragging a video progress bar to adjust playback of exciting segments, this invention offers more convenient switching operations and ensures that the target video segment being played is complete, unaffected by the accuracy of the user's dragging operation.

[0066] This embodiment of the video skip point generation method includes: acquiring the video to be segmented and its video data; the video data includes at least: dialogue information, video description, and user behavior data; segmenting the video to be segmented into video segments of different program types based on the dialogue information and video description; determining the user's points of interest in the video segments based on the user behavior data; generating skip points for the video to be segmented based on the dialogue information, video description, and points of interest; the skip points represent the client's response to an adjustment command, jumping the playback progress of the video to be segmented to the start time of the target video segment. Segmenting video segments by program type solves the problem that video segments obtained by scene division cannot adapt to different program types; determining the user's points of interest based on user behavior data and segmenting video segments according to these points of interest ensures the quality of the content played at the generated skip points.

[0067] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0068] Reference Figure 3 The diagram illustrates the structure of a video skip point generation device provided in an embodiment of the present invention, which may specifically include the following: The video to be segmented acquisition module 301 is used to acquire the video to be segmented and the video data of the video to be segmented; the video data includes at least: dialogue information, video description and user behavior data; The video segmentation module 302 is used to segment the video to be segmented into video segments of different program types according to the dialogue information and the video description; The point of interest generation module 303 is used to determine the user's points of interest in the video segment based on the user behavior data. The jump point generation module 304 is used to generate jump points for the video to be segmented based on the dialogue information, the video description, and the points of interest; the jump point indicates that the client responds to the adjustment command and jumps the playback progress of the video to be segmented to the start time point of the target video segment.

[0069] In one embodiment of the present invention, the video segmentation module 302 includes: The program type determination submodule is used to determine the program type contained in the video to be segmented based on the dialogue information and the video description; The segmentation strategy matching submodule is used to match the corresponding video segmentation strategy according to the program type; The video segmentation module is used to segment the video to be segmented into multiple video segments according to the video segmentation strategy and the program type.

[0070] In one embodiment of the present invention, the program type includes at least: performance programs and talk shows; the video segmentation strategy includes: a first video segmentation strategy and a second video segmentation strategy; The first segmentation unit is used to extract the first starting time point of each performance in the video to be segmented when the video to be segmented contains performance programs, using a first video segmentation strategy; and to segment the video to be segmented into video segments containing performance programs according to the first starting time point. The second segmentation unit is used to extract the second starting time point of each talk show in the video to be segmented according to the second video segmentation strategy when the video to be segmented contains talk shows, and input the second starting time point, the dialogue information and the video synopsis into a pre-trained large model, and determine the structured segments of the talk shows through the large model; and segment the video to be segmented into video segments containing the talk shows according to the second starting time point and the structured segments.

[0071] In one embodiment of the present invention, the user behavior data includes at least: a playback curve and a video speed adjustment point; the video speed adjustment point represents the maximum value point at which the user adjusts the video to be segmented from a first speed to a second speed, wherein the first speed is greater than the second speed; the point of interest generation module 303 includes: The maximum point determination submodule is used to determine the maximum point of the playback curve based on the playback curve. The point of interest determination submodule is used to merge the maximum replay point with the video speed adjustment point to obtain the point of interest.

[0072] In one embodiment of the present invention, the jump view point generation module 304 includes: The target video segment segmentation module is used to perform secondary segmentation of the video segment based on the points of interest and a preset attention duration threshold to obtain multiple target video segments. The jump view point generation submodule is used to determine the jump view points of the video to be segmented based on the dialogue information, the video description, and the start time of the target video segment.

[0073] In one embodiment of the present invention, the target video segment segmentation module includes: The first segmentation unit for the target video segment is configured to perform secondary segmentation of the video segment based on the interest point when the segment length of the video segment is greater than the preset attention duration threshold, thereby obtaining the target video segment; the target video segment contains the interest point. The second segmentation unit for the target video segment is used to identify the video segment as the target video segment when the segment duration of the video segment is less than or equal to the preset attention duration threshold.

[0074] In one embodiment of the present invention, the device further includes: The video script generation module is used to generate video scripts for the target video clip based on the dialogue information and video description contained in the target video clip; The adjustment instruction acquisition module is used to acquire the user's adjustment instructions for the preset switching area on the playback page of the video to be segmented; The display module is used to, upon receiving the adjustment instruction, jump the playback progress of the video to be segmented to the nearest jump point, play the target video segment, and display the video text.

[0075] For the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple; relevant details can be found in the descriptions of the method embodiments. This invention also provides an electronic device, such as... Figure 4 As shown, it includes a processor 401, a communication interface 402, a memory 403, and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404. Memory 403 is used to store computer programs; When processor 401 executes the program stored in memory 403, it performs the following steps: Obtain the video to be segmented and the user behavior data of the video to be segmented; The video to be segmented is divided into first video segments of different program types according to the video scene; The first video segment is divided into two segments based on the program type of the first video segment to obtain the second video segment. The user's points of interest in the video to be segmented are determined based on the user behavior data. The second video segment is segmented according to the points of interest to obtain target video segments that meet the preset attention duration threshold and jump points of the target video segments; the jump points are used to switch the target video segments according to the user's adjustment instructions on the playback page of the video to be segmented.

[0076] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0077] The communication interface is used for communication between the aforementioned terminal and other devices.

[0078] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0079] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0080] like Figure 5As shown, this embodiment of the invention also provides a computer-readable storage medium 501, which stores instructions that, when run on a computer, cause the computer to execute any of the video jump point generation methods described in the above embodiments.

[0081] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the video jump point generation methods described in the above embodiments.

[0082] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0083] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0084] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0085] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A method for generating video skip points, characterized in that, The method includes: Obtain the video to be segmented and its video data; the video data includes at least: dialogue information, video description, and user behavior data; Based on the dialogue information and the video description, the video to be segmented is divided into video segments of different program types; Determine the user's points of interest in the video clip based on the user behavior data; The jump points of the video to be segmented are generated based on the dialogue information, the video description, and the points of interest; the jump points indicate that the client responds to the adjustment command and jumps the playback progress of the video to be segmented to the start time point of the target video segment.

2. The method according to claim 1, characterized in that, The step of segmenting the video to be segmented into video segments of different program types based on the dialogue information and the video description includes: The program types contained in the video to be segmented are determined based on the dialogue information and the video description. Match the corresponding video segmentation strategy according to the program type; The video to be segmented is divided into multiple video segments according to the video segmentation strategy and the program type.

3. The method according to claim 2, characterized in that, The program types include at least: performance programs and talk shows; the video segmentation strategies include: a first video segmentation strategy and a second video segmentation strategy; The step of dividing the video to be divided into multiple video segments according to the video segmentation strategy and the program type includes: If the video to be segmented contains performance programs, a first video segmentation strategy is used to extract the first starting time point of each performance program in the video to be segmented; and the video to be segmented is segmented into video segments containing performance programs based on the first starting time point. If the video to be segmented contains talk shows, the second starting time point of each talk show in the video to be segmented is extracted according to the second video segmentation strategy. The second starting time point, the dialogue information and the video synopsis are then input into a pre-trained large model to determine the structured segments of the talk shows. Based on the second starting time point and the structured segments, the video to be segmented is segmented into video segments containing the talk shows.

4. The method according to claim 1, characterized in that, The user behavior data includes at least: a playback curve and a video speed adjustment point; the video speed adjustment point represents the maximum value point at which the user adjusts the video to be segmented from the first speed to the second speed, wherein the first speed is greater than the second speed; Determining the user's interest in the video clip based on the user behavior data includes: The maximum value point of the playback curve is determined based on the playback curve. The point of interest is obtained by fusing the maximum value point during playback with the video speed adjustment point.

5. The method according to claim 1, characterized in that, The step of generating jump points for the video to be segmented based on the dialogue information, the video description, and the points of interest includes: The video segment is further segmented based on the points of interest and a preset attention duration threshold to obtain multiple target video segments. The jump view points of the video to be segmented are determined based on the dialogue information, the video description, and the start time of the target video segment.

6. The method according to claim 5, characterized in that, The video segment is further segmented based on the points of interest and a preset attention duration threshold to obtain multiple target video segments, including: If the duration of the video segment exceeds the preset attention duration threshold, the video segment is further segmented based on the interest points to obtain a target video segment; the target video segment contains the interest points. If the duration of the video segment is less than or equal to the preset attention duration threshold, the video segment is used as the target video segment.

7. The method according to claim 1, characterized in that, The method further includes: Generate video script for the target video segment based on the dialogue information and video description contained in the target video segment; Upon receiving the adjustment instruction, the playback progress of the video to be segmented is redirected to the nearest skip point, the target video segment is played, and the video text is displayed.

8. A video skip point generation device, characterized in that, The device includes: The video to be segmented acquisition module is used to acquire the video to be segmented and the video data of the video to be segmented; the video data includes at least: dialogue information, video description and user behavior data; The video segmentation module is used to segment the video to be segmented into video segments of different program types based on the dialogue information and the video description. An interest point generation module is used to determine the user's interest points in the video clip based on the user behavior data; The jump point generation module is used to generate jump points for the video to be segmented based on the dialogue information, the video description, and the points of interest; the jump point indicates that the client responds to the adjustment command and jumps the playback progress of the video to be segmented to the start time point of the target video segment.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.