Video watching point generation method and device, video playing method and device based on video watching point, electronic equipment and storage medium

By automatically segmenting videos and generating highlights using a large model, this technology solves the problem of strong subjectivity in human judgment of highlights in videos, achieving more accurate segmentation of highlights and improving user experience.

CN120916016APending Publication Date: 2025-11-07BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511067921.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing technologies, the determination of whether a video highlight is exciting or not mainly relies on human judgment, which is highly subjective and makes it difficult to accurately identify the highlights of a video.

Method used

The system utilizes a large model to obtain single-episode summary information, text, and scene information of the target video. Based on preset highlights conditions, it automatically segments the video and generates highlights information. By matching text and scene information with the large model, it generates highlights of exciting video clips, enabling exciting plot jumps during video playback.

Benefits of technology

It achieves more accurate and objective segmentation of exciting video highlights, allowing users to quickly understand video clips, satisfy their viewing needs for engaging storylines, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120916016A_ABST
    Figure CN120916016A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video watching point generation method and device, a video playing method and device based on a video watching point, electronic equipment and a storage medium, and relates to the technical field of Internet application, and the video watching point generation method comprises the steps: obtaining single set summary information of a target video, and text information and scene information in the target video; segmenting the target video at a plurality of time points by using the large model according to a preset wonderful watching point condition to obtain a plurality of first video clips; for each first video clip, matching the single-set summary information with text information and scene information in the first video clip by using a large model to generate first watching point information, so as to respond to the first triggering operation when the target video is played on the playing page; and determining a first video clip closest to the current time point based on the time point corresponding to the first watching point information, skipping to the time point corresponding to the first video clip from the target video, and playing. According to the embodiment of the invention, more accurate wonderful video watching points can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of Internet application, in particular to a video highlight generation method and device, a video playing method and device based on video highlight, an electronic device and a storage medium. BACKGROUND

[0002] In order to adjust the viewing rhythm, users of a video website usually adjust the viewing progress by dragging the progress bar, but it is difficult to smoothly jump to the content to be watched by manual operation, in order to solve this problem, the video website has launched a video highlight function, such as displaying time point positions and corresponding plots on the video progress bar, so that users can jump to watch according to the video highlight.

[0003] However, the video highlight in the prior art can only be simply divided according to the plot changes, if the division according to whether the plot is exciting or not is to be realized, artificial participation is often needed, the subjectivity is too strong, and it is difficult to obtain relatively accurate exciting video highlights. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a video highlight generation method and device, a video playing method and device based on video highlight, an electronic device and a storage medium, so as to obtain more accurate exciting video highlights, and the specific technical solutions are as follows:

[0005] In a first aspect, the embodiments of the present application provide a video highlight generation method, comprising:

[0006] obtaining single-set abstract information of a target video, and text information and scene information appearing in the target video, wherein the single-set abstract information comprises a plot outline of the target video;

[0007] based on the text information and the scene information, using a large model to divide the target video at multiple time points according to a preset exciting highlight condition, to obtain multiple first video clips;

[0008] for each first video clip, using the large model to match the single-set abstract information with the text information and the scene information appearing in the first video clip, to generate first highlight information of the first video clip, so that when the target video is played on a playing page, in response to a first trigger operation, the closest first video clip to the current time point is determined based on the first highlight information, and the target video is jumped to the time point corresponding to the first video clip and played.

[0009] In a second aspect, the embodiments of the present application provide a video playing method based on video highlight, comprising:

[0010] playing a target video on a playing page, the target video being associated with first highlight information, the first highlight information being obtained based on the video highlight generation method in any of the above and corresponding to a first video segment, the first video segment being obtained by dividing the target video according to a preset highlight condition based on a large model, and first highlight information of each first video segment being sequentially ordered according to time points in single-set abstract information of the target video;

[0011] in response to a first trigger operation, determining a first video segment closest to a current time point based on the first highlight information;

[0012] jumping to a time point corresponding to the first video segment in the target video and playing.

[0013] In a third aspect, an embodiment of the present application provides a video highlight generation device, comprising:

[0014] an information acquisition module configured to acquire single-set abstract information of a target video, and text information and scene information appearing in the target video, wherein the single-set abstract information comprises a plot outline of the target video;

[0015] a video segmentation module configured to segment the target video at multiple time points based on the text information and the scene information and according to a preset highlight condition by using a large model, to obtain multiple first video segments;

[0016] a highlight generation module configured to, for each first video segment, match the single-set abstract information with text information and scene information appearing in the first video segment by using a large model, to generate first highlight information of the first video segment, so that, when the target video is played on a playing page, in response to a first trigger operation, a first video segment closest to a current time point is determined based on the first highlight information, and the target video is jumped to a time point corresponding to the first video segment and played.

[0017] In a fourth aspect, an embodiment of the present application provides a video playing device based on video highlights, comprising:

[0018] a video playing module configured to play a target video on a playing page, the target video being associated with first highlight information, the first highlight information being obtained based on the video highlight generation device described above and corresponding to a first video segment, the first video segment being obtained by dividing the target video according to a preset highlight condition based on a large model, and first highlight information of each first video segment being sequentially ordered according to time points in single-set abstract information of the target video;

[0019] An operation response module is configured to, in response to a first trigger operation, determine a first video segment closest to a current time point based on the first highlight information.

[0020] A play jump module is configured to jump the target video to a time point corresponding to the first video segment and play the target video.

[0021] In another aspect of the present application, an electronic device is provided, which comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory are in communication with each other through the communication bus.

[0022] The memory is configured to store a computer program.

[0023] The processor is configured to execute the program stored in the memory, thereby implementing the method steps of any of the above aspects.

[0024] In another aspect of the present application, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the method steps of any of the above aspects.

[0025] In another aspect of the present application, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the video highlight generation method and the video play method based on video highlights.

[0026] In another aspect of the present application, a computer program product is provided, which comprises instructions, and when the computer program product is run on a computer, the computer is caused to execute the video highlight generation method and the video play method based on video highlights.

[0027] The video highlight generation method provided in this invention obtains single-episode summary information of the target video, as well as text information and scene information appearing in the target video. Based on the text information and scene information, a large model is used to segment the target video at multiple time points according to preset highlight conditions, resulting in multiple first video segments. Because the large model itself has a massive database and knowledge base, and possesses complex reasoning capabilities, it can understand complex instructions such as preset highlight conditions and reason about these conditions based on massive amounts of data, thereby obtaining first video segments divided according to highlight conditions. Compared to manual evaluation of highlight levels, segmenting videos according to preset highlight conditions using a large model can yield more accurate and objective highlight data. For each first video segment, a large model is used to match the summary information of the single episode with the text and scene information appearing in the first video segment to generate the first highlight information of the first video segment. The single episode summary information includes a plot summary of the target video, while the highlight information represents a plot summary of the first video segment. This allows for a concise summary of the first video segment, enabling viewers to quickly understand its content based on the highlights. Each first highlight information is associated with the target video so that when the target video is played on the playback page, in response to the first trigger operation, the user can jump from the target video to the corresponding time point of the first video segment and start playing it. This allows for seamless transitions between exciting plot points, satisfying users' viewing needs for engaging storylines and improving the user experience. This interactive adjustment allows for a faster viewing pace and a more focused viewing experience. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0029] Figure 1 This is a flowchart illustrating a video highlight generation method provided in an embodiment of the present invention;

[0030] Figure 2 This is one possible implementation of step S102 provided in the embodiments of the present invention;

[0031] Figure 3 This is one possible implementation of step S203 provided in the embodiments of the present invention;

[0032] Figure 4 This is one possible implementation of step S103 provided in the embodiments of the present invention;

[0033] Figure 5 This is one possible implementation of step S404 provided in the embodiments of the present invention;

[0034] Figure 6 This is a flowchart illustrating a video highlight generation method provided in an embodiment of the present invention.

[0035] Figure 7-1 This is a schematic diagram of the first process of a video playback method based on video highlights provided in an embodiment of the present invention;

[0036] Figure 7-2 This is a schematic diagram illustrating a plot jump example according to an embodiment of the present invention;

[0037] Figure 7-3 This is a schematic diagram illustrating an example of a playback page according to an embodiment of the present invention;

[0038] Figure 8 This is a schematic diagram of a second process for a video playback method based on video highlights provided in an embodiment of the present invention;

[0039] Figure 9 This is a schematic diagram of a video highlight generation device provided in an embodiment of the present invention;

[0040] Figure 10 This is a schematic diagram of the structure of a video playback device based on video highlights, provided in an embodiment of the present invention.

[0041] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0042] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0043] Because existing technologies can only categorize video highlights based on simple plot changes, and categorizing them based on the excitement of the plot often requires human intervention, which is too subjective and makes it difficult to obtain relatively accurate highlights, this invention provides a video highlight generation method and apparatus, a video playback method and apparatus based on video highlights, an electronic device, and a storage medium.

[0044] The video highlight generation method and the video playback method based on video highlights provided in this embodiment of the invention can be applied to any electronic device with video playback capabilities, such as computers, TVs, and mobile terminals.

[0045] The video highlight generation method provided by the present invention will be described in detail below through specific embodiments.

[0046] See Figure 1 , Figure 1 This is a schematic diagram of a first step in the video highlight generation method according to an embodiment of the present invention, including:

[0047] In step S101, single-set abstract information of a target video is obtained, as well as text information and scene information appearing in the target video.

[0048] The single-set abstract information includes a plot outline of the target video.

[0049] The target video needs to be divided into multiple exciting video clips and video highlights corresponding to each video clip are generated. The video clip can be understood as an exciting plot unit. The single-set abstract information is a plot outline of the target video, which includes detailed plot development content in the target video. For example, when the target video is a certain episode of a series video, the single-set abstract information includes a plot introduction of the episode and a single-episode introduction. Exemplarily, the single-set abstract information can be text information provided by an uploader of the target video and uploaded together with the target video, or public information summarized by a third-party website or a third-party author according to the content of the target video, which is obtained by an information grabbing algorithm.

[0050] The text information appearing in the target video includes dialogues between characters, monologues of characters, and voice-overs appearing in the video, and also includes text information appearing on the video screen, such as introduction text of scenes and characters in the screen.

[0051] In an example, since the text information in the video usually has no punctuation marks, after obtaining the text information of the target video, a large model is used to perform semantic analysis on the text information to realize sentence breaking processing of the video text information and add punctuation marks to the text information. Exemplarily, the text information can be input into the large model, and an instruction of semantic analysis and adding punctuation marks is input into the large model, so that the large model outputs the text information after sentence breaking with punctuation marks.

[0052] The scene information refers to a scene appearing in the video, such as indoor, outdoor, street, park, etc. Under the same scene, even if the scene information corresponding to the screen shot from different angles is still the same scene information. Exemplarily, the scene information of the target video can be obtained by a video scene detection algorithm, an image recognition algorithm, etc.

[0053] The large model refers to an artificial intelligence large model, including a large language model, a visual large model, a multi-modal large model, and a basic scientific large model, etc., which is trained by a large amount of data sets, has a large number of parameters, and has the ability of understanding, generating, and complex reasoning according to instructions.

[0054] In step S102, based on the text information and the scene information, a large model is used to divide the target video at multiple time points according to preset exciting highlight conditions, to obtain multiple first video clips.

[0055] The preset highlight condition indicates whether the current video content is exciting, is a pre-set, large model understandable, and a judgment condition for whether the plot is exciting. Specifically, it can include whether it is related to the main character; whether it can promote the main story line, such as plot turning points, plot suspense, and major discoveries within the plot; whether it is the content or picture that the audience likes, and can also include more detailed judgment conditions that the large model can understand, such as whether it includes a comic scene, a comedy laugh point, an intense fighting scene, a suspenseful plot, and the like.

[0056] Exemplarily, the preset highlight condition is exemplarily illustrated, and the exciting plot includes, but is not limited to, the following types: climax scene, emotional explosion, surprise twist, action scene, funny moment, key revelation, and romantic moment. Among them, the climax scene refers to the moment when the plot reaches the highest tension, usually when the main character faces a major challenge or makes a critical decision; the emotional explosion refers to the emotional dialogue or conflict between characters, which can be a moving confession, a fierce argument, or a moving reunion; the surprise twist refers to an unexpected reversal or major change in the plot, which surprises or surprises the audience; the action scene refers to an exciting fight, chase or other high-difficulty action scene, with excellent visual effects and tension; the funny moment refers to a humorous plot or dialogue that makes the audience laugh out loud; the key revelation refers to the revelation of important information or truth in the plot, which has a major impact on the overall story; and the romantic moment refers to the moment of romantic and sweet interaction between lovers, which brings warmth and beauty to the audience. When the plot in the video segment is the above-mentioned exciting plot, the video segment meets the highlight condition.

[0057] The preset highlight condition is input into the large model as an instruction to enable the large model to learn the exciting plot. Since the large model itself has a vast database and knowledge base, and has complex reasoning capabilities, it can understand complex instructions such as the exciting plot condition (e.g., related to the main character, promoting the main story line, etc.) and can reason whether the audience likes this type of condition based on the vast database.

[0058] In one example, the large model can be pre-trained according to the preset highlight condition to enable the large model to have more accurate recognition of the current preset highlight condition when it already has reasoning capabilities for the highlight condition.

[0059] Exemplarily, the target video can be input into the large model, the preset highlight condition and the video segmentation instruction can be input into the large model, and the large model can be obtained according to the highlight condition, the time point of each first video segment, and the first video segment segmented according to the time point.

[0060] The target video is divided according to preset highlight conditions, so that each video segment divided is a video segment meeting the highlight conditions, that is, the first video segment obtained is a highlight video segment of the target video divided according to the highlights, and each first video segment corresponds to a highlight plot unit.

[0061] In step S103, for each first video segment, the single-set abstract information is matched with text information and scene information appearing in the first video segment by using a large model to generate first highlight information of the first video segment, so that when the target video is played on a playing page, in response to a first trigger operation, the first highlight information is used to determine a first video segment closest to a current time point and jump to a time point corresponding to the first video segment and play from the time point.

[0062] Specifically, the single-set abstract information includes a plot outline of the target video, and the text information and scene information in each video segment are matched, and the matching result is content of the single-set abstract information corresponding to each video segment. The content corresponding to each video segment is extracted from the single-set abstract information to obtain highlight information of the video segment. The highlight information of the video segment represents the plot outline of the video segment, which can be summarized in a relatively short text information to summarize the first video segment, so that a user can quickly understand the content of the video segment according to the video highlight.

[0063] For example, the text information and scene information corresponding to the first video segment and the single-set abstract information of the target video can be input into a large model, a highlight generation instruction is input into the large model, and the large model generates highlight information of the first video segment according to the context information.

[0064] After the first highlight information of each video segment is extracted, the first highlight information can be associated with the target video, for example, a jump point is set in the target video according to a time point of the first video segment corresponding to the first highlight information, or the first highlight information is bound to a predetermined operation to jump between time points based on the predetermined operation. Thus, in the video playing process, after receiving the predetermined operation, the next time point can be jumped to and the video data can be played from the time point, so that when a user is not interested in the current plot, the next highlight plot can be jumped to, the user's demand for watching the highlight plot is met, and the user experience is improved.

[0065] As can be seen from the above, the video highlight generation method provided by the embodiment of the present application obtains the single-set abstract information of the target video, and the text information and the scene information appearing in the target video; based on the text information and the scene information, the target video is divided at multiple time points by using a large model according to preset highlight conditions, and multiple first video segments are obtained; since the large model itself has a mass database and a knowledge base, and has complex reasoning capability, the first video segments divided according to the highlight conditions can be obtained by understanding the complex instructions of the preset highlight conditions and reasoning the preset highlight conditions according to the mass data. Compared with manual judgment of highlight degree, the highlight video points can be obtained more accurately and objectively by dividing according to the preset highlight conditions by using the large model. For each first video segment, the single-set abstract information and the text information and the scene information appearing in the first video segment are matched by using the large model to generate first highlight information of the first video segment, the single-set abstract information includes a plot outline of the target video, and the highlight information represents a plot outline of the first video segment, so that the first video segment can be summarized by using relatively short text information, and a watching user can quickly understand the content of the video segment according to the video highlight. The first highlight information is associated with the target video, so that when the target video is played on a playing page, the first video segment corresponding to a time point is jumped to and played in response to a first trigger operation, the highlight scenes can be played by jumping between the highlight scenes, the watching demand of a user for the highlight scenes is met, and the user experience is improved. This interactive adjustment can make the watching rhythm of the user faster and the watching content of the user more focused.

[0066] In an embodiment of the present application, as shown in Figure 2 The step S102 based on the text information and the scene information, utilizes a large model to divide the target video at multiple time points according to preset highlight conditions, and obtains multiple first video segments, including:

[0067] Step S201, according to the scene information of the target video, obtaining a scene change corresponding transition time point in the target video;

[0068] Step S202, based on the transition time point, the target video is divided into multiple second video segments;

[0069] Step S203, in the case where the time length of the second video segment exceeds a first preset time length threshold, the second video segment is divided by using a large model according to preset highlight conditions, and each first video segment is obtained.

[0070] The transition indicates a change of scene in the video, for example, the scene in the picture changes from A to B. It should be noted that only different roles of the perspective and lens switching are not transitions, and only the change of the overall scene is considered as a transition. The transition time point is the time point of the scene change.

[0071] According to the picture information of the target video, a plurality of scenes appearing in the target video are determined, the picture information refers to the picture data corresponding to each frame of the target video, the scene information of the target video is determined according to the statistics of the picture information, and then the time point of the change of the scene information is obtained, and the transition time point is obtained.

[0072] The target video is segmented based on the transition time point, for example, the video is divided into S0, S1, … Si, Si+1, … and a plurality of second video segments according to the change of the scene. If the duration of the second video segment does not exceed the preset duration threshold, it indicates that the video segment is short enough, and further segmentation may cause the video to be too fragmented and difficult to generate effective video highlights. For example, the target video can be input into a large model, and the scene change instruction is input into the large model, so that the large model segments the target video according to the transition time point to obtain the second video segment.

[0073] In the case where the duration of the second video segment exceeds the first preset duration threshold (pre-set according to actual needs, for example, 1 minute, 3 minutes, 5 minutes, etc.), the large model is used to segment the second video segment according to the preset highlight condition to obtain the first video segment segmented according to the highlight.

[0074] For example, the second video segment can be input into a large model, and the preset highlight condition and the video segmentation instruction are input into the large model to obtain the first video segment segmented by the large model according to the highlight condition, the time point of each first video segment, and the first video segment segmented according to the time point.

[0075] As can be seen from the above, the video highlight generation method provided by the embodiment of the application first segments the target video according to the scene change to obtain the second video segment, and then segments according to the preset highlight condition in the case where the duration of the second video segment exceeds the first preset duration threshold, to obtain the first video segment segmented according to the highlight, thereby reducing the problem of repeated or invalid video highlights caused by too short video segment duration.

[0076] In one embodiment of the application, as shown in Figure 3 The above step S203 segments the second video segment according to the preset highlight condition by using the large model in the case where the duration of the second video segment exceeds the first preset duration threshold, to obtain each first video segment, including:

[0077] Step S301, obtaining text information and picture information appearing in the second video segment;

[0078] Step S302, generating video description information of the second video segment according to the text information and the picture information by using a large model;

[0079] The video description information includes a summary of the second video segment.

[0080] Step S303, analyzing the text information, the picture information and the video description information by using a large model, and segmenting the second video segment according to a preset highlight condition to obtain each first video segment.

[0081] The time length of the first video segment is less than the first preset time length threshold.

[0082] The text information appearing in the second video segment includes dialogue information appearing in the second video segment, text information in the picture, etc., which can be obtained from the text information corresponding to the target video according to the time point of the second video segment in the target video, for example, the time point of scene Si is ts-te, and the content of the text information corresponding to the target video ts-te is obtained. The picture information refers to the picture data of the second video segment, which can be, for example, one frame of image data per second in the video segment, for example, one frame per second in the ts-te segment to obtain an image.

[0083] In the case where the time length of the second video segment exceeds the first preset time length threshold, the large model analyzes the semantic information in the text information and the picture content corresponding to the picture information to generate the video description information of the video segment, which includes a summary of the second video segment, or a description of the key seconds of the second video segment, for example, describing that characters A and B have D events in scene C.

[0084] For example, the text information and the picture information of the second video segment can be input into the large model, and a video description instruction can be input into the large model to make the large model generate the video description information of the video segment according to the text information and the picture information.

[0085] The large model infers the highlight division time point of the second video segment according to the text information, the picture information and the video description information of the second video segment, and divides each second video segment into a first video segment, so that the time length of the first video segment is less than the first preset time length threshold.

[0086] Exemplarily, the text information, picture information and video description information of the second video segment can be input into the large model, an instruction according to a preset highlight condition is input into the large model, and the large model is caused to infer a highlight division time point of the second video segment and output the first video segment.

[0087] As can be seen from the above, the video highlight generation method provided in the embodiment of the application, in the case that the time length of the second video segment exceeds the first preset time length threshold, uses a large model to analyze the text information, picture and video description information of the second video segment, divides the second video segment into the first video segment according to the highlight condition, obtains the first video segment further divided according to the highlight, and improves the accuracy and objectivity of the highlight division.

[0088] In an embodiment of the application, the above method further comprises:

[0089] In the case that the time length of the first video segment is lower than a second preset time length threshold, the text information of the first video segment and the single-set abstract information of the target video are analyzed by using a large model, at least two continuous first video segments that meet a preset aggregation condition are aggregated, and a first video segment with a time length lower than the first preset time length threshold and not lower than the second preset time length threshold is obtained.

[0090] The second preset time length threshold is less than the first preset time length threshold.

[0091] The second preset time length threshold is set in advance according to actual needs, for example, 30 seconds, 10 seconds, etc. When the time length is lower than the second preset time length threshold, it indicates that the video segment division is too fragmented, and meaningless video highlight division is generated. Therefore, the large model is used to merge the video segments based on the context semantic scene. Specifically, the text information and single-set abstract information of the first video segment with a too short time length are obtained, the large model is used for semantic analysis, the continuous first video segments that meet the preset aggregation condition are merged in time sequence, and the first video segment with a time length lower than the first preset time length threshold and not lower than the second preset time length threshold is obtained.

[0092] Specifically, the preset aggregation condition can be set according to actual needs, or can be completely left to the large model to judge, for example, whether two video segments are different scenes but in the same plot or promote the development of the same plot.

[0093] Exemplarily, the first video segment with a time length lower than the second preset time length threshold and its continuous previous video segment or next video segment can be input into the large model, a merging instruction according to the preset aggregation condition is input into the large model, and the large model outputs the merged first video segment.

[0094] For example, according to the dialogue of Si and Si+1 and the single-set abstract of the target video, the large model judges whether Si and Si+1 can be further merged. If the large model judges that Si and Si+1 can be merged, and the time after Si and Si+1 are merged is less than a first preset time length threshold (for example, 3 minutes), Si and Si+1 are merged, and a plurality of scenes (first video segments) M0, M1, …, Mi are obtained after a round of merging.

[0095] As can be seen from the above, the video highlight generation method provided in the embodiment of the application analyzes the first video segment with a time length less than a second preset time length threshold, and merges the continuous video segments that meet the preset aggregation condition, thereby reducing the problem of meaningless video highlight division caused by too fragmented video segment division.

[0096] In an embodiment of the application, as shown in Figure 4 The step S103 uses a large model to match the single-set abstract information with the text information and scene information appearing in the first video segment, to generate the first highlight information of the first video segment, including:

[0097] In the step S401, in a case where the number of text information appearing in the first video segment is less than a preset text quantity threshold, a multi-modal abstract of the first video segment is generated based on the text information and picture information of the first video segment.

[0098] In the step S402, a large model is used to match the multi-modal abstract with the single-set abstract information of the target video, to extract the abstract information matching successfully from the single-set abstract information, and obtain the plot abstract information of the first video segment.

[0099] In the step S403, in a case where the number of text information appearing in the first video segment is not less than a preset text quantity threshold, a large model is used to match the text information of the first video segment with the single-set abstract information, to extract the abstract information matching successfully from the single-set abstract information, and obtain the plot abstract information of the first video segment.

[0100] In the step S404, the first highlight information of the first video segment is generated according to the plot abstract information.

[0101] The plot abstract, i.e. the abstract information, of each first video segment represents the plot summary of the first video segment, which can enable a watching user to quickly understand the plot content of the video segment.

[0102] In a case where the number of text information appearing in the first video segment is lower than the preset text quantity threshold, for example, less than three sentences, five sentences, etc., it is indicated that the plot content of the video segment cannot be accurately determined from the text information alone, and then a multi-modal summary of the video segment is generated based on the text information and the picture information of the video segment by using the large model.

[0103] For example, the picture information of the video segment can include multiple shots, each shot can correspond to a piece of multi-modal summary, and the multi-modal summary of the video segment is a combination of the multi-modal summaries of each shot. The multi-modal summary can include text information, video information, image information, etc., and is obtained according to the actual situation of the video segment, for example, can include video shot frame extraction of several seconds, tens of seconds, etc.

[0104] The multi-modal summary of the video segment is matched with the single-set summary of the target video by using the large model to determine which part of the plot in the video segment is in the single-set summary, and the matching result obtained is the part corresponding to the multi-modal summary of the video segment in the single-set summary. The part matching successfully in the single-set summary is extracted to obtain the plot summary information of the first video segment.

[0105] For example, the single-set summary includes that the characters a and b experience events 1, 2 and 3 together, and the multi-modal summary of the video segment indicates that the characters a and b are experiencing event 1. After matching the two, the matching result obtained is the content corresponding to event 1, and the part of content corresponding to event 1 is extracted from the single-set summary as the plot summary information of the video segment.

[0106] For example, the multi-modal summary of the first video segment and the single-set summary of the target video can be input into the large model, a content matching instruction is input into the large model, and the large model outputs the plot summary information of the first video segment.

[0107] In a case where the number of text information appearing in the first video segment is not lower than the preset text quantity threshold, it is indicated that there is more text information (for example, more dialogues) in the first video segment, and the plot content of the video segment can be more accurately determined from the text information alone. Then, the text information is matched with the single-set summary information by using the large model to determine which part of the text information of the video segment corresponds to the single-set summary of the target video, and the matching result obtained is the part of content corresponding to the video segment in the single-set summary. This part of content is extracted to generate the plot summary information of the first video segment.

[0108] For example, the text information of the first video segment and the single-set summary of the target video can be input into the large model, a content matching instruction is input into the large model, and the large model outputs the plot summary information of the first video segment.

[0109] The episode summary information is a piece of text information indicating a content outline of the first video segment. The episode summary information can be used as the highlight information of the first video segment.

[0110] For example, if there is more redundant content in the episode summary, the episode summary can also be highlighted to obtain the highlight information after simplification. This part can also be realized by a large model. The episode summary is input into the large model, and a highlight instruction is input to make the large model output the simplified highlight information.

[0111] As can be seen from the above, the video highlight generation method provided by the embodiment of the application generates a multi-modal summary of a video segment by a large model in the case that the number of text information in the first video segment is less than the preset text quantity threshold, matches and extracts the content outline of the first video segment, generates highlight information based on this, and further improves the accuracy of the highlight information. In the case that the number of text information in the first video segment is not less than the preset text quantity threshold, the episode summary information is extracted according to the text information and the single-set summary information, and then the highlight information is generated, thereby reducing the inference calculation amount of the large model and improving the efficiency of the highlight information generation.

[0112] In an embodiment of the application, as shown in Figure 5 The step S404 generates the first highlight information of the first video segment according to the episode summary information, and includes the following steps.

[0113] In step S501, for each first video segment, whether the episode summary information of the first video segment causes a spoiler for the continuous next video segment is identified by a large model based on the episode summary information of the first video segment, the episode summary information of the continuous next video segment and the text information.

[0114] In step S502, in the case of causing a spoiler, the spoiler content is removed from the episode summary information of the first video segment to obtain the episode summary information after removing the spoiler content.

[0115] In step S503, whether the episode summary information of the first video segment after removing the spoiler content has duplicate content with the second highlight information of the continuous previous video segment is analyzed by a large model.

[0116] In step S504, in the case of having duplicate content, the duplicate content of the episode summary information is removed to obtain the first highlight information of the first video segment.

[0117] The large model analyzes the summary information of the first video segment to determine whether it reveals the content of the next video segment. If the large model determines that the summary information of the current video segment reveals the content of the next video segment, the part of the content that causes the spoiler is removed to obtain the plot summary information after removing the spoiler content.

[0118] For example, the summary information of the first video segment, the plot summary information and the text information of the next video segment can be input into the large model, and an instruction to analyze whether it is a spoiler can be input into the large model, so that the large model outputs whether it is a spoiler and the content that causes the spoiler.

[0119] The large model analyzes the plot summary information of the first video segment after removing the spoiler content to determine whether it has duplicate information with the historical highlight information of the previous video segment. If there is duplicate information, the duplicate content is removed to obtain the first highlight information of the first video segment. Specifically, when analyzing the highlight information of the current video segment, the plot summary of the current video segment after removing the spoiler content and the historical highlight information (the highlight information of the previous video segment before the current video segment or all video segments before that) can be input into the large model. The large model can also generate the story line of the target video. Based on the highlight information of each video segment and the story line, it is determined whether there is duplicate content. If there is, the duplicate content is removed to generate the final highlight information.

[0120] For example, the story line can be obtained by the large model analyzing the plot summary / highlight information of each video segment in the target video. It is a more concise version of the plot summary, which matches each video segment and refers to the plot development of the target video.

[0121] A specific example is provided below:

[0122] The plot summary information of video segment A is: Jiang Yi, Jiang Mu, and Lu Er and others take the elevator in the hospital together. The elevator is crowded with seven or eight people. Jiang Mu reminds others to wait, which may be to let more people into the elevator. When the elevator reaches the first floor, Lu Yan reminds everyone that they have arrived, and they are ready to move the patient's bed. In the crowded elevator, Jiang Yi and Lu Er stand very close to each other, and their hands almost touch each other, showing the subtle changes in their relationship.

[0123] The story line is: the relationship between Jiang Yi and Lu Er has changed subtly.

[0124] Historical highlights:

[0125] Highlight 1: Lu Er dreams of the past between him and Jiang Yi;

[0126] Highlight 2: Lu Er handles multiple emergency surgeries in a row;

[0127] Point 3: Jiang Yi refused father's surgery due to risk;

[0128] If it is judged according to the story line, the plot summary information of the video segment A and the historical points that there is no repeated content in the plot summary information, no modification is needed, and the point information of the video segment A is directly generated according to the plot summary information.

[0129] For example, the plot summary of the video segment A includes that the event a occurs between the role 1 and the role 2, includes more detailed event content, and achieves the result b, and the story line of the video segment A can be that the role 1 and the role 2 achieve the result b. The historical points include that the event c occurs between the role 1 and the role 2 in the video segment B, and achieves the result d, and the story line of the video segment B can be that the role 1 and the role 2 achieve the result d. The large model can analyze whether there is repeated content only according to the historical points and the story line of each video segment, and if there is, the repeated content is removed to generate the point information.

[0130] For example, the plot summary information of the current video segment after the elimination of the spoiler content can be input into the large model, an instruction of analyzing whether there is repetition is input into the large model, the large model outputs the result of whether there is repetition and the repeated content.

[0131] As can be seen from the above, the video point generation method provided by the embodiment of the present application judges whether there is spoiler or repeated information in the summary information of the first video segment, and further improves the effectiveness of the point information.

[0132] In an embodiment of the present application, the above method further comprises:

[0133] In the case that the large model analyzes that the first point information of the first video segment is meaningless point, the first video segment is merged with the continuous previous video segment, wherein the meaningless point includes point information without effective plot.

[0134] In the case that the large model analyzes that the first point information of the first video segment is meaningless point, for example, there is no effective plot in the video segment itself, and the large model forcibly extracts the point, the video segment is merged with the previous video segment.

[0135] For example, the first point information of the first video segment can be input into the large model, and an instruction of analyzing whether it is meaningful is input into the large model, so that the large model outputs the result of whether it is meaningful.

[0136] As can be seen from the above, the video point generation method provided by the embodiment of the present application merges the video segment corresponding to the meaningless point, and reduces the redundant information caused by too many invalid points.

[0137] In one embodiment of the present application, a flowchart example diagram of a video highlight generation method is also provided. Based on a large model, highlight points are divided and highlight scripts are generated. Based on dialogues, transition points, and multi-modal descriptions, video streams are divided to construct plot segments. Based on single-episode plot summaries, dialogues, and video descriptions of the video stream, a plot summary is generated. The use of storylines and summaries to generate highlight scripts can effectively identify highlight points, eliminate watered-down plots, and provide a new method for users to watch long videos.

[0138] As shown in the example, Figure 6 A flowchart example diagram of a video highlight generation method is provided. The original dialogues (punctuation-free text), transition points (scene switching points), plot summaries (single-episode summary information, single-episode plot synopsis), and multi-modal video descriptions (picture descriptions) of the target video are obtained. The dialogues are processed intelligently, and punctuation is added to the original dialogues. The scenes are merged and divided, wherein scenes with the same semantics are merged based on context semantics, and scenes that are too long are divided based on video descriptions and pictures. Based on multi-modal data, a plot summary is generated for each video segment, wherein a large model is used to intelligently match multi-modal descriptions to generate a plot summary, or a plot summary is generated based on dialogues and multi-modal descriptions. A large model is used to intelligently filter meaningless highlights. Under the guidance of the story line (narrative context) of the target video, plot highlights are generated based on the plot summary and the story line, and meaningless highlights are intelligently filtered to obtain highlight points and plot highlights of exciting video segments.

[0139] Referring to Figure 7-1 , a flowchart of a first video playback method based on video highlights is shown.

[0140] Step S701, playing a target video on a playback page;

[0141] The target video is associated with first highlight information, the first highlight information is obtained based on the video highlight generation method described in any of the above embodiments, and corresponds to a first video segment. The first video segment is obtained by dividing the target video based on a large model according to a preset exciting highlight condition. The first highlight information of each first video segment is sorted in chronological order in the single-episode summary information of the target video.

[0142] After the target video is associated with the first highlight information, if the server subsequently receives a playback request from the client, it can send the video data of the target video to the client. The video data is parsed and rendered on the client, and the target video is played on the playback page.

[0143] As Figure 7-2As shown, the target video is associated with first highlight information, for example, the first video clip of plot 1 is located at the start of 22:30 to 23:54 of the video, titled B scolds A, the first video clip of plot 2 is located at the start of 23:54 to 24:39 of the video, titled A discovers the spirit stone.

[0144] The target video is associated with first highlight information, and the generation method of the first highlight information is described above and will not be repeated here.

[0145] When playing the target video, the terminal device such as a mobile phone or a tablet computer can start the skip play function under a set model. In an optional embodiment of the present application, when the terminal device is in a horizontal screen mode, at least one jump operation area is set in the play page, and the first trigger operation includes a sliding operation received in the jump operation area. The jump operation area is an operation area for triggering the jump play of the first video clip. The position of the area can be set based on requirements, for example, set at the left side, right side, lower side or upper and lower sides of the play page. The terminal device can also be switched to a horizontal screen mode to display prompt information to prompt the viewing user to trigger the jump play of the first video clip.

[0146] In one example, operation prompt information can be displayed on the play page, such as Figure 7-3 In the example shown, jump operation areas are set on the left and right of the play page, which can prompt "scan and jump to watch plots up and down". A brightness operation area is set in the left 2 area, which prompts "slide up and down to adjust brightness", and a volume operation area is set in the right 2 area, which prompts "slide up and down to adjust volume". Thus, the user can perform operation processing in the corresponding area to adjust the brightness, volume and jump between different plots during the process of watching the video.

[0147] In step S702, in response to the first trigger operation, the first video clip closest to the current time point is determined based on the first highlight information.

[0148] During the process of watching the target video in the play page, if the user wants to switch to other plot clips, the user can perform the first trigger operation, and the corresponding client can receive the first trigger operation, such as up and down sliding operation, gesture operation, etc. In response to the first trigger operation, the time point of the current play of the video is determined, and the first video clip closest to the time point of the current play is determined in the first highlight information as the target highlight plot unit.

[0149] As shown in Figure 7-2 In response to the first trigger operation such as a sliding operation, the time point of plot 1 is jumped to the time point of plot 2.

[0150] The first trigger operation includes a forward jump operation and / or a backward jump operation. In response to the first trigger operation, a first video clip closest to a current time point is determined based on the first highlight information, including:

[0151] If a forward jump operation is received, a first video clip corresponding to a time point before the current playing time point is searched in each first highlight information as a target highlight unit. For example, based on the received downward sliding operation, a first video clip corresponding to a time point before the current playing time point is searched in each first highlight information as a target highlight unit, and playing starts from the target highlight unit.

[0152] If a backward jump operation is received, a first video clip corresponding to a time point after the current playing time point is searched in each first highlight information as a target highlight unit. For example, based on the received upward sliding operation, a first video clip corresponding to a time point after the current playing time point is searched in each first highlight information as a target highlight unit, and playing starts from the target highlight unit.

[0153] In step S703, the target video is jumped to the time point corresponding to the first video clip and played.

[0154] The video data corresponding to the first video clip (target highlight unit) is obtained, the video data is parsed and rendered, and the target video is jumped to the video data rendered by the first video clip and played.

[0155] In the embodiment of the application, the target video can jump between the first video clips. In response to a second trigger operation, playing jumps between the first video clips. The second trigger operation is an operation for triggering jumping between the first video clips. The first trigger operation and the second trigger operation can be the same operation or different operations, and the embodiment of the application does not limit this. After receiving the second trigger operation, in response to the second trigger operation, playing jumps between the first video clips. For example, if the second trigger operation is a forward jump operation (or playing the previous plot), a previous highlight unit of the current first video clip is determined, and playing jumps to the previous first video clip. For example, if the second trigger operation is a backward jump operation (or playing the next plot), a next first video clip of the current first video clip is determined, and playing jumps to the next first video clip. Thus, the video data jumps between the video clips, improving the user experience of watching the video.

[0156] The above-mentioned manner not only retains the complete narrative of the long video, but also enables the user to enjoy the viewing rhythm similar to the short video, effectively improving the user experience and platform retention rate.

[0157] Referring to Figure 8 , a flowchart of a second video playing method based on video highlights of the application is shown.

[0158] Step S701, playing a target video in a playing page.

[0159] The target video is associated with first highlight information.

[0160] Step S702, in response to a first trigger operation, determining a first video segment closest to a current time point based on the first highlight information.

[0161] Step S703, jumping the target video to a time point corresponding to the first video segment and playing.

[0162] Step S704, displaying a title of the first video segment in the playing page.

[0163] If the server receives a playing request from the client, it can send video data of the target video to the client, which parses and renders the video data and plays the target video in the playing page. The client receives a first trigger operation, such as up and down sliding operations, gesture operations, etc. In response to the first trigger operation, if a forward jump operation is received, the first highlight information is searched for a first video segment before the current playing time point, which is taken as a target highlight unit. If a backward jump operation is received, the first highlight information is searched for a first video segment after the current playing time point, which is taken as a target highlight unit.

[0164] The video data corresponding to the first video segment is obtained, the video data is parsed and rendered, the target video is jumped to the video data rendered from the first video segment to start playing, and the title is displayed in the playing page. As shown in Figure 7-2 , in response to the first trigger operation, the playing is started from the time point of plot 1 to the time point of plot 2, and the title of plot 2 is displayed in the playing page: character A discovers the spirit stone.

[0165] At present, some users are used to watching short videos, and when watching long videos, they are prone to distraction, often choosing to directly exit when encountering uninteresting segments, resulting in a continuous decline in user retention rate of long video platforms. Among them, short videos refer to videos with a total duration of not more than a first time threshold, such as videos with a total duration of not more than 3, 4, 5, 6 minutes, etc., while long videos refer to videos with a total duration of more than a second time threshold, such as videos with a total duration of more than 20, 25, 30, 40, 45 minutes, etc.

[0166] The embodiment of the present application can solve the problem of the split between long video watching experience and user habits, and by introducing the advantages of short videos into the long video watching scene, the intelligent regulation of content rhythm is realized, the complete narrative of long videos is maintained, and the rapid consumption demand of users for climax content is met. Specifically, the present application uses the above processing steps to analyze the structure of long videos, deeply understands and evaluates the video content, sets screening rules and standards according to different plot types (such as suspense, costume idol drama, fantasy, modern idol drama, etc.), realizes intelligent identification and extraction of the first video segment corresponding to the wonderful plot, uses a multi-stage processing method, refines the plot of the video segment, and reorganizes the long video into a video segment with high attraction. Through simple and intuitive sliding gestures, sliding and other trigger operations, users can easily jump between these wonderful segments and enjoy the viewing rhythm similar to short dramas. This innovative processing method not only retains the complete narrative of long videos, but also allows users to enjoy the viewing rhythm similar to short videos, effectively improving user experience and platform retention rate.

[0167] Referring to Figure 9 The embodiment of the present application also provides a structural diagram of a video highlight generation device, which comprises:

[0168] An information acquisition module 901 is configured to acquire single-set abstract information of a target video, and text information and scene information appearing in the target video, wherein the single-set abstract information comprises a plot outline of the target video.

[0169] A video segmentation module 902 is configured to segment the target video at multiple time points based on the text information and scene information, and using a large model according to preset wonderful highlight conditions, to obtain multiple first video segments.

[0170] A highlight generation module 903 is configured to, for each first video segment, match the single-set abstract information with the text information and scene information appearing in the first video segment by using a large model, to generate first highlight information of the first video segment, so that when the target video is played on a playing page, in response to a first trigger operation, the nearest first video segment to a current time point is determined based on the first highlight information, and the target video is jumped to the first video segment corresponding to the time point and played.

[0171] In an embodiment of the present application, the video segmentation module 902 comprises:

[0172] A transition time point acquisition sub-module is configured to acquire transition time points corresponding to scene changes in the target video according to the scene information of the target video.

[0173] The first video segmentation module is configured to segment the target video into a plurality of second video segments based on the transition time point.

[0174] The second video segmentation module is configured to, in a case where a time length of the second video segment exceeds a first preset time length threshold, segment the second video segment according to a preset highlight condition by using a large model to obtain the first video segments.

[0175] In an embodiment of the present application, the second video segmentation module is specifically configured to:

[0176] obtain text information and picture information appearing in the second video segment;

[0177] generate video description information of the second video segment according to the text information and the picture information by using a large model, wherein the video description information includes a summary of the second video segment;

[0178] analyze the text information, the picture information and the video description information according to a preset highlight condition by using a large model to segment the second video segment to obtain the first video segments, wherein a time length of the first video segment is less than the first preset time length threshold.

[0179] In an embodiment of the present application, the device further comprises:

[0180] The video aggregation module is configured to, in a case where a time length of the first video segment is less than a second preset time length threshold, analyze text information of the first video segment and single-set abstract information of the target video by using a large model to aggregate at least two continuous first video segments satisfying a preset aggregation condition to obtain a first video segment with a time length less than the first preset time length threshold and not less than the second preset time length threshold, wherein the second preset time length threshold is less than the first preset time length threshold.

[0181] In an embodiment of the present application, the highlight generation module 903 comprises:

[0182] The multi-modal abstract generation submodule is configured to, for each first video segment, in a case where a quantity of text information appearing in the first video segment is less than a preset text quantity threshold, generate a multi-modal abstract of the first video segment based on the text information and picture information of the first video segment.

[0183] The single-set abstract information generation submodule is configured to match the multi-modal abstract with single-set abstract information of the target video by using a large model, extract abstract information matching successfully from the single-set abstract information to obtain plot abstract information of the first video segment.

[0184] The plot abstract information generation submodule is configured to, in a case where the quantity of text information appearing in the first video segment is not less than a preset text quantity threshold, match the text information of the first video segment with the single-set abstract information by using a large model, extract abstract information with which the matching is successful from the single-set abstract information, and obtain plot abstract information of the first video segment.

[0185] The highlight information generation submodule is configured to generate first highlight information of the first video segment according to the plot abstract information.

[0186] In an embodiment of the present application, the highlight information generation submodule is specifically configured to:

[0187] For each first video segment, based on the plot abstract information of the first video segment, the plot abstract information of the continuous next video segment, and the text information, a large model is used to identify whether the plot abstract information of the first video segment causes a spoiler for the continuous next video segment.

[0188] In a case where a spoiler is caused, spoiler content is removed from the plot abstract information, and plot abstract information after the spoiler content is removed is obtained.

[0189] A large model is used to analyze whether the plot abstract information of the first video segment after the spoiler content is removed has repeated content with second highlight information of the continuous previous video segment.

[0190] In a case where repeated content exists, the repeated content of the plot abstract information is removed, and first highlight information of the first video segment is obtained.

[0191] In an embodiment of the present application, the device further comprises:

[0192] The video merging module is configured to, in a case where a large model analyzes that the first highlight information of the first video segment is meaningless highlight, merge the first video segment with the continuous previous video segment, wherein the meaningless highlight includes highlight information without effective plot.

[0193] Referring to Figure 10 The embodiment of the present application further provides a structural schematic diagram of a video playing device based on video highlights, which comprises:

[0194] The video playback module 1001 is used to play a target video on the playback page. The target video is associated with first highlight information. The first highlight information is obtained based on the video highlight generation device and corresponds to a first video segment. The first video segment is obtained by dividing the target video according to preset highlights conditions based on a large model. The first highlight information of each first video segment is sorted in the single-episode summary information of the target video according to the time point order.

[0195] Operation response module 1002 is used to respond to the first trigger operation and determine the first video segment closest to the current time point based on the first viewpoint information;

[0196] The playback jump module 1003 is used to jump the target video to the corresponding time point of the first video segment and play it.

[0197] This invention also provides an electronic device, such as... Figure 11 As shown, it includes a processor 1101, a communication interface 1102, a memory 1103, and a communication bus 1104. The processor 1101, communication interface 1102, and memory 1103 communicate with each other via the communication bus 1104.

[0198] Memory 1103 is used to store computer programs;

[0199] The processor 1101, when executing the program stored in the memory 1103, implements any of the video highlight generation methods described above.

[0200] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0201] The communication interface is used for communication between the aforementioned terminal and other devices.

[0202] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0203] The processor described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; or can be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0204] In yet another embodiment provided by the present application, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the video highlight generation method according to any one of the above embodiments.

[0205] In yet another embodiment provided by the present application, a computer program product is provided, and the computer program product includes instructions. When the computer program product is executed on a computer, the computer is caused to perform the video highlight generation method according to any one of the above embodiments.

[0206] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof. When implemented by software, the implementation can be in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the implementation produces the processes or functions according to the embodiments of the present application. The computer can be a general purpose computer, a special purpose computer, a computer network or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0207] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0208] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0209] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A video highlight generation method, characterized by, The method comprises: obtaining single-set abstract information of a target video, and text information and scene information appearing in the target video, wherein the single-set abstract information comprises a plot outline of the target video; based on the text information and scene information, using a large model to split the target video at multiple time points according to preset highlight conditions, to obtain multiple first video segments; for each first video segment, using a large model to match the single-set abstract information with the text information and scene information appearing in the first video segment, to generate first highlight information of the first video segment, so that when the target video is played on a playing page, in response to a first trigger operation, based on a time point corresponding to the first highlight information, a first video segment closest to a current time point is determined, and the target video is jumped to a time point corresponding to the first video segment and played.

2. The method of claim 1, wherein, The method further comprises: based on the text information and scene information, using a large model to split the target video at multiple time points according to preset highlight conditions, to obtain multiple first video segments, comprising: obtaining transition time points corresponding to scene changes in the target video according to the scene information of the target video; based on the transition time points, splitting the target video into multiple second video segments; 3. The method of claim 2, wherein, in a case where the time length of the second video segment exceeds a first preset time length threshold, using a large model to split the second video segment according to preset highlight conditions, to obtain each first video segment. The method further comprises: in a case where the time length of the second video segment exceeds a first preset time length threshold, using a large model to split the second video segment according to preset highlight conditions, to obtain each first video segment, comprising: obtaining text information and picture information appearing in the second video segment; 4. The method of claim 1, wherein, using a large model to generate video description information of the second video segment according to the text information and picture information, wherein the video description information comprises a plot outline of the second video segment; using a large model to analyze the text information, the picture information, and the video description information, and splitting the second video segment according to preset highlight conditions to obtain each first video segment, wherein the time length of the first video segment is lower than the first preset time length threshold.

5. The method of claim 1, wherein, The method further comprises: in a case where the time length of the first video segment is lower than a second preset time length threshold, using a large model to analyze the text information of the first video segment and the single-set abstract information of the target video, and aggregating at least two continuous first video segments that satisfy a preset aggregation condition to obtain a first video segment with a time length lower than the first preset time length threshold and not lower than the second preset time length threshold, wherein the second preset time length threshold is lower than the first preset time length threshold. The method further comprises: for each first video segment, using a large model to match the single-set abstract information with the text information and scene information appearing in the first video segment, to generate first highlight information of the first video segment, comprising: In a case where the number of text information appearing in the first video segment is less than a preset text quantity threshold, a multi-modal summary of the first video segment is generated based on the text information and picture information of the first video segment; The multi-modal summary is matched with single-set summary information of the target video by using a large model, and matching successful summary information is extracted from the single-set summary information to obtain plot summary information of the first video segment; In a case where the number of text information appearing in the first video segment is not less than a preset text quantity threshold, the text information of the first video segment is matched with the single-set summary information by using a large model, and matching successful summary information is extracted from the single-set summary information to obtain plot summary information of the first video segment; First highlight information of the first video segment is generated according to the plot summary information.

6. The method of claim 5, wherein, The first highlight information of the first video segment is generated according to the plot summary information, including: For each first video segment, whether the plot summary information of the first video segment causes a spoiler for a continuous next video segment is identified by using a large model based on the plot summary information of the first video segment, the plot summary information and text information of the continuous next video segment; In a case where a spoiler is caused, spoiler content is removed from the plot summary information of the first video segment to obtain plot summary information after the spoiler content is removed; Whether there is repeated content between the plot summary information of the first video segment after the spoiler content is removed and second highlight information of a continuous previous video segment is analyzed by using a large model; In a case where there is repeated content, the repeated content of the plot summary information is removed to obtain first highlight information of the first video segment.

7. The method of claim 1, wherein, The method further includes: In a case where the first highlight information of the first video segment is meaningless highlight information by large model analysis, the first video segment is merged with a continuous previous video segment, wherein the meaningless highlight information includes highlight information without effective plot.

8. A video playback method based on video highlights, characterized in that, Including: A target video is played on a playing page, wherein the target video is associated with first highlight information, the first highlight information is obtained based on the video highlight generation method in any one of claims 1-7, and corresponds to a first video segment, the first video segment is obtained based on a large model dividing the target video according to a preset highlight condition, and first highlight information of each first video segment is sorted in a time point order in single-set summary information of the target video; In response to a first trigger operation, a first video segment closest to a current time point is determined based on the first highlight information; The target video is jumped to a time point corresponding to the first video segment and played.

9. A video highlight generation apparatus, characterized by comprising: Including: An information acquisition module is configured to acquire single-set summary information of a target video and text information and scene information appearing in the target video, wherein the single-set summary information includes a plot outline of the target video; The video cutting module is configured to cut the target video at multiple time points based on the text information and the scene information and according to preset highlight conditions by using a large model, so as to obtain multiple first video clips. The highlight generation module is configured to, for each first video clip, match the single-set abstract information with text information and scene information appearing in the first video clip by using a large model, and generate first highlight information of the first video clip, so that, when the target video is played on a playing page, the first highlight information is used to determine a first video clip closest to a current time point in response to a first trigger operation, and the target video is jumped to a time point corresponding to the first video clip and played. 10.A video playing device based on video highlights, characterized in that, The video playing module is configured to play a target video on a playing page, the target video being associated with first highlight information, the first highlight information being obtained based on the video highlight generation apparatus in claim 9 and corresponding to a first video clip, the first video clip being obtained by dividing the target video based on a large model and according to preset highlight conditions, and the first highlight information of each first video clip being sorted in a time point sequence in single-set abstract information of the target video. The operation response module is configured to determine a first video clip closest to a current time point in response to a first trigger operation based on the first highlight information. The playing jump module is configured to jump the target video to a time point corresponding to the first video clip and play the target video. The apparatus includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory are in communication with each other via the communication bus.

11. An electronic device, comprising: The memory is configured to store a computer program. The processor is configured to execute the program stored in the memory, and implement the method steps in any one of claims 1-8. The computer program stored in the computer readable storage medium is executed by the processor, and implements the method steps in any one of claims 1-8.

12. A computer-readable storage medium, characterized in that, ​