Video segmentation method, device and equipment, storage medium, program product and video understanding model training method

By iteratively processing video keyframes and using a video understanding model, combined with accumulated state information and descriptive text, the problem of inaccurate video segmentation in existing technologies is solved, achieving higher segmentation accuracy and logical coherence.

CN122223628APending Publication Date: 2026-06-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2026-04-09
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing technologies lack the ability to understand contextual semantics based on plot facts, resulting in inaccurate video segmentation boundaries, overly fragmented segmentation results, and low fault tolerance, making it difficult to meet the needs of deep plot segmentation in complex video scenarios.

Method used

By iteratively processing the keyframes of the target video, combining a preset time window and a video understanding model, the video is segmented using accumulated state information and descriptive text, and the segmentation points are dynamically adjusted to improve accuracy and logical coherence.

Benefits of technology

It significantly improves the accuracy of video segmentation results and the logical coherence and semantic integrity of each video segment, effectively avoiding erroneous segmentation caused by sudden changes in local scenes or complex editing techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223628A_ABST
    Figure CN122223628A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video segmentation method, device, equipment, storage medium, program product and training method of video understanding model, the method comprises: obtaining the current key frame sequence and cumulative state information of the target video in the current time window; based on the current key frame sequence and the cumulative state information, determining the third description text of each current key frame in the current key frame sequence, the second identifier of the cumulative segmentation point, and the fourth description text of each video segment obtained after segmentation; based on the third description text, the second identifier and the fourth description text, updating the cumulative state information to obtain the cumulative state information of the next time window of the current time window; when the iteration processing of all key frames in the target video is completed, based on the cumulative state information obtained after the iteration processing, the target video is segmented to obtain the video segmentation result of the target video. Through the present application, the accuracy of video segmentation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the Internet field, and includes, but is not limited to, a video segmentation method, apparatus, device, storage medium, program product, and training method for a video understanding model. Background Technology

[0002] With the rapid development of internet video technology, major video platforms generate massive amounts of video content such as movies, TV dramas, and variety shows every day. A typical episode of a mainstream TV drama is between 40 and 60 minutes long, and the entire series comprises dozens of episodes, totaling dozens of hours of video content. To facilitate users' quick understanding of the core content of videos, achieve accurate video segment recommendations, and enable intelligent retrieval, there is an urgent need for automated story segmentation and plot summarization of video content.

[0003] Among related technologies, one type of method performs video segmentation based on visual features. This involves extracting features such as color histograms, edge features, or motion vectors from video frames and detecting scene boundaries based on differences or similarities between adjacent video frames, thereby dividing the video into multiple video segments. Another type of method performs video segmentation based on deep learning scene information. This involves using scene recognition models to extract background scene features, image content features, or scene embedding features from video frames, and then segmenting the video content based on feature clustering results, similarity results, or classification results. A third type of method performs video segmentation based on specific rules. This involves pre-setting program structure rules for specific types of videos and integrating information such as intros, titles, host features, camera changes, silence points, cut points, and audio cycle changes to construct an event sequence in chronological order, and then determining the start and end positions of video segments according to predetermined rules. In addition, there are methods for video segmentation based on multimodal information fusion. These methods typically analyze audio and image data in the video separately. For example, they use an audio model to obtain audio segmentation results and a scene recognition model to obtain image segmentation results. Then, they merge the audio segmentation results and the image segmentation results to obtain the final video segmentation result.

[0004] However, the relevant technologies lack the ability to understand contextual semantics based on plot facts, resulting in inaccurate segmentation boundaries, overly fragmented segmentation results, and low fault tolerance in complex video scenes, making it difficult to meet the actual needs of current video deep plot segmentation. Summary of the Invention

[0005] This application provides a video segmentation method, apparatus, device, storage medium, program product, and training method for a video understanding model, which can improve the accuracy of video segmentation.

[0006] The technical solution of this application embodiment is implemented as follows: This application provides a video segmentation method, the method comprising: performing the following iterative processing on keyframes of a target video according to a preset time window: obtaining the current keyframe sequence and cumulative state information of the target video within the current time window; the cumulative state information including: a first description text for each historical keyframe in a historical keyframe sequence prior to the current time window, a first identifier of a historical segmentation point corresponding to the historical keyframe sequence, and a second description text for each video segment obtained after segmenting the historical keyframe sequence according to the historical segmentation point; and determining a third description text for each current keyframe in the current keyframe sequence based on the current keyframe sequence, the first description text, the first identifier, and the second description text. The system comprises: a second identifier of the cumulative segmentation point corresponding to the cumulative keyframe sequence; and a fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation point; the cumulative keyframe sequence includes the current keyframe sequence and the historical keyframe sequence; the cumulative state information is updated based on the third descriptive text, the second identifier, and the fourth descriptive text to obtain the cumulative state information of the next time window of the current time window; the current time window and the next time window partially overlap in the time dimension; when the iterative processing of all keyframes in the target video is completed, the target video is segmented based on the cumulative state information obtained after the iterative processing to obtain the video segmentation result of the target video.

[0007] This application provides a video segmentation device, comprising: a first iterative processing module, configured to perform the following iterative processing on keyframes of a target video according to a preset time window: acquiring the current keyframe sequence and cumulative state information of the target video within the current time window; the cumulative state information including: a first description text for each historical keyframe in a historical keyframe sequence preceding the current time window, a first identifier of a historical segmentation point corresponding to the historical keyframe sequence, and a second description text for each video segment obtained after segmenting the historical keyframe sequence according to the historical segmentation point; and determining a third description text for each current keyframe in the current keyframe sequence based on the current keyframe sequence, the first description text, the first identifier, and the second description text. The video segmentation module is configured to perform video segmentation on the target video based on the cumulative segmentation information obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation information, the second identifier of the cumulative segmentation point corresponding to the cumulative segmentation point, and the fourth description text of each video segment obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation point; the cumulative keyframe sequence includes the current keyframe sequence and the historical keyframe sequence; the cumulative state information is updated based on the third description text, the second identifier, and the fourth description text to obtain the cumulative state information of the next time window of the current time window; the current time window and the next time window partially overlap in the time dimension; the video segmentation module is configured to perform video segmentation on the target video based on the cumulative state information obtained after the iterative processing of all keyframes in the target video to obtain the video segmentation result of the target video.

[0008] In the above scheme, the first iterative processing module is further configured to: obtain object reference information corresponding to the target video, and text information corresponding to each current keyframe in the current keyframe sequence; fill the object reference information, the text information, the first descriptive text, the first identifier, and the second descriptive text into a preset prompt word template to obtain the target prompt word; input the current keyframe sequence and the target prompt word into a pre-trained video understanding model; and determine the third descriptive text, the second identifier, and the fourth descriptive text through the video understanding model.

[0009] In the above scheme, the first iterative processing module is further configured to: generate an initial result for the current keyframe sequence using the video understanding model; the initial result includes an initial frame description text for each current keyframe in the current keyframe sequence, an initial identifier for the initial cumulative segmentation point corresponding to the cumulative keyframe sequence, and an initial segment description text for each first video segment obtained after segmenting the cumulative keyframe sequence according to the initial cumulative segmentation point; calibrate the initial frame description text based on the preset calibration text in the target prompt word to obtain the third description text; calibrate the initial identifier to obtain the second identifier; and calibrate the initial segment description text based on the second identifier to obtain the fourth description text.

[0010] In the above scheme, the first iterative processing module is further configured to: for each current keyframe in the current keyframe sequence, obtain the current text corresponding to the current keyframe from the target prompt word; when the current text includes a first object identifier of the target object in the current keyframe, determine a second object identifier of the target object in the current keyframe from the initial frame description text; when the first object identifier is different from the second object identifier, replace the second object identifier in the initial frame description text with the first object identifier to obtain the third description text; when the current text does not include any first object identifier, or when the first object identifier is the same as the second object identifier, determine the initial frame description text as the third description text.

[0011] In the above scheme, the first iterative processing module is further configured to: for any initial cumulative segmentation point, based on the third descriptive text and the first descriptive text in the target prompt word, determine the semantic coherence between the preceding and following video segments corresponding to the initial cumulative segmentation point; when the semantic coherence indicates semantic coherence between the preceding and following video segments, delete the initial identifier of the initial cumulative segmentation point; when the semantic coherence indicates semantic incoherence between the preceding and following video segments, determine the initial identifier of the initial cumulative segmentation point as the second identifier.

[0012] In the above scheme, the first iterative processing module is further configured to: segment the accumulated keyframe sequence according to the initial accumulated segmentation point corresponding to the second identifier to obtain multiple second video segments; for each second video segment, when the second video segment includes multiple first video segments, semantically merge the initial segment description texts of the multiple first video segments to obtain a fourth description text of the second video segment; when the second video segment includes one first video segment, determine the initial segment description text of the first video segment as the fourth description text of the second video segment.

[0013] In the above scheme, the first iterative processing module is further configured to: concatenate the third description text after the first description text in the cumulative state information to obtain a merged description text; determine the merged description text as the first description text in the cumulative state information of the next time window, determine the second identifier as the first identifier in the cumulative state information of the next time window, and determine the fourth description text as the second description text in the cumulative state information of the next time window.

[0014] In the above scheme, the video segmentation module is further configured to: obtain an identifier sequence corresponding to the identifiers of all segmentation points of the target video from the accumulated state information obtained after iterative processing; perform video segmentation on the target video sequentially based on the identifiers in the identifier sequence to obtain multiple target video segments; and determine the multiple target video segments as the video segmentation result of the target video.

[0015] In the above scheme, the video segmentation device further includes a keyframe acquisition module; the keyframe acquisition module is used to: extract frames from the target video according to a preset sampling rate to obtain an initial video frame sequence; determine the scene embedding features and text information of each initial video frame in the initial video frame sequence; perform scene segmentation on the initial video frame sequence based on the scene embedding features to obtain multiple video scenes; and perform deduplication processing on the initial video frames in each video scene based on the scene embedding features and the text information to obtain the keyframe sequence of the target video.

[0016] In the above scheme, the keyframe acquisition module is further configured to: determine the feature similarity between the scene embedding features of each of the two adjacent initial video frames in the initial video frame sequence; normalize the feature similarity to obtain a normalized feature similarity; when the normalized feature similarity is less than a preset first feature similarity threshold, determine the later initial video frame among the two adjacent initial video frames as the segmentation point; and perform segmentation on the initial video frame sequence based on the segmentation point to obtain multiple video segments.

[0017] In the above scheme, the keyframe acquisition module is further configured to: extract a reference video frame sequence of a preset duration from the target video; determine the reference feature similarity between the scene embedding features of each two adjacent reference video frames in the reference video frame sequence; determine the mean similarity and standard deviation of the multiple reference feature similarities; for each feature similarity of each two adjacent initial video frames in the initial video frame sequence, determine the difference between the feature similarity and the mean similarity; and determine the ratio of the difference to the standard deviation of similarity as the normalized feature similarity.

[0018] In the above scheme, the keyframe acquisition module is further configured to: for each video segment, when any two adjacent initial video frames in the video segment meet at least one of the following preset conditions, determine the text quality of the text information of each of the two adjacent initial video frames; wherein, the preset conditions are: the feature similarity between the scene embedding features of each of the two adjacent initial video frames in the video segment is greater than a preset second feature similarity threshold; the text similarity between the text information of each of the two adjacent initial video frames is greater than a preset text similarity threshold; when the text quality corresponding to the earlier initial video frame among the two adjacent initial video frames is greater than or equal to a preset quality threshold, retain the earlier initial video frame; when the text quality corresponding to the earlier initial video frame is less than the quality threshold, retain the initial video frame corresponding to the largest text quality among the two adjacent initial video frames; sort the retained initial video frames according to time order to obtain the keyframe sequence of the target video.

[0019] This application provides a training method for a video understanding model. The method includes: performing the following iterative processing on sample keyframes of sample videos in a training sample set according to a preset training time window: obtaining the current keyframe sequence of the sample video within the current training time window, sample cumulative state information, and target output ground truth; the sample cumulative state information includes: a sample first description text for each sample historical keyframe in the sample historical keyframe sequence located before the current training time window, a sample first identifier for the sample historical segmentation point corresponding to the sample historical keyframe sequence, and each sample keyframe obtained after segmentation according to the sample first identifier. The sample video segment's second description text; the target output truth value includes: a frame description truth value sequence for the sample keyframes of the sample video, a truth value sequence for the actual segmentation points of the sample video, and a segment description truth value sequence for the sample video segment obtained after segmenting the sample video according to the actual segmentation points; noise injection processing is performed on the sample first description text to obtain noisy first description text; noise injection processing is performed on the sample first identifier to obtain noisy first identifier; noise injection processing is performed on the sample second description text to obtain noisy second description text; based on the current keyframe sequence of the sample, the noisy first description text... The noisy first identifier and the noisy second description text are used by a video understanding model to determine a prediction result. The prediction result includes: a predicted third description text for each current keyframe in the current keyframe sequence of the sample, a predicted second identifier for the cumulative segmentation point corresponding to the cumulative keyframe sequence of the sample, and a predicted fourth description text for each video segment obtained after segmenting the cumulative keyframe sequence of the sample according to the cumulative segmentation point. The cumulative keyframe sequence of the sample includes the current keyframe sequence of the sample and the historical keyframe sequence of the sample. The prediction is based on the predicted third description text and the frame description ground truth. The loss is calculated using the sequence, the predicted second identifier and the identifier ground truth sequence, and the predicted fourth description text and the segment description ground truth sequence to obtain a multi-task joint loss value; based on the predicted third description text, the predicted second identifier and the predicted fourth description text, the sample cumulative state information is updated to obtain the sample cumulative state information for the next training time window of the current training time window; the current training time window and the next training time window partially overlap in the time dimension; based on the multi-task joint loss value, the network parameters of the video understanding model to be trained are updated to obtain the trained video understanding model.

[0020] In the above scheme, the step of calculating the first loss based on the predicted third descriptive text and the target frame descriptive ground truth to obtain the keyframe descriptive loss value includes: merging the one-hot encodings of the positions of each word in the target frame descriptive ground truth in the preset dictionary into a first one-hot encoding sequence; obtaining the first prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting and outputting the predicted third descriptive text; determining the first cross-entropy loss value based on the first prediction probability in the first prediction probability distribution matrix and the first one-hot encoding label in the first one-hot encoding sequence; and determining the first cross-entropy loss value as the keyframe descriptive loss value.

[0021] In the above scheme, the step of calculating the second loss based on the predicted second identifier and the ground truth value of the target identifier to obtain the video segmentation loss value includes: calculating the cross-entropy loss based on the predicted second identifier and the ground truth value of the target identifier for each sample cumulative segmentation point corresponding to the sample cumulative keyframe sequence to obtain the cross-entropy loss value; and aggregating the cross-entropy loss values ​​of all sample cumulative segmentation points corresponding to the sample cumulative keyframe sequence to obtain the video segmentation loss value.

[0022] In the above scheme, the step of calculating the third loss based on the predicted fourth descriptive text and the target fragment description ground value to obtain the fragment description loss value includes: merging the one-hot encodings of the positions of each word in the target fragment description ground value in the preset dictionary into a second one-hot encoding sequence; obtaining the second prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting and outputting the predicted fourth descriptive text; determining the second cross-entropy loss value based on the second prediction probability in the second prediction probability distribution matrix and the second one-hot encoding label in the second one-hot encoding sequence; and determining the second cross-entropy loss value as the fragment description loss value.

[0023] In the above scheme, the step of calculating the fourth loss based on the prediction result and the target output ground truth to obtain the global loss value includes: according to a preset output format, summarizing the target frame description ground truth value, the target identifier ground truth value, and the target segment description ground truth value in the target output ground truth value into a target output word sequence; merging the one-hot encodings of the positions of each word in the target output word sequence in the preset dictionary into a third one-hot encoding sequence; obtaining the third prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting and outputting the prediction result; determining the third cross-entropy loss value based on the third prediction probability in the third prediction probability distribution matrix and the third one-hot encoding label in the third one-hot encoding sequence; and determining the third cross-entropy loss value as the global loss value.

[0024] This application provides a training apparatus for a video understanding model, comprising: a second iterative processing module, configured to perform the following iterative processing on sample keyframes of sample videos in a training sample set according to a preset training time window: acquiring the current keyframe sequence of the sample video within the current training time window, sample cumulative state information, and target output ground truth; the sample cumulative state information includes: a first description text of each sample historical keyframe in the sample historical keyframe sequence located before the current training time window, a first identifier of the sample historical segmentation point corresponding to the sample historical keyframe sequence, and each sample keyframe obtained after segmentation according to the first identifier. The sample video segment's sample second description text; the target output truth value includes: a frame description truth value sequence for the sample keyframes of the sample video, a truth value sequence for the identifiers of the sample video's actual segmentation points, and a segment description truth value sequence for the sample video segment obtained after segmenting the sample video according to the actual segmentation points; noise injection processing is applied to the sample first description text to obtain noisy first description text; noise injection processing is applied to the sample first identifier to obtain noisy first identifier; noise injection processing is applied to the sample second description text to obtain noisy second description text; based on the current keyframe sequence of the sample, the noisy first description text, and the noisy second description text, the target output truth value includes: a frame description truth value sequence for the sample keyframes of the sample video, a truth value sequence for the identifiers of the sample video's actual segmentation points, and a segment description truth value sequence for the sample video segment obtained after segmenting the sample video according to the actual segmentation points; noise injection processing is applied to the sample first description text to obtain noisy first description text; noise injection processing is applied to the sample second description text to obtain noisy second description text; based on the sample current keyframe sequence, the noisy first description text, and the noisy second description text, the target output truth value sequence includes: a frame description truth value sequence for the sample keyframes of the sample video, a truth value sequence for the identifiers of the sample video's actual segmentation points, and a truth value sequence for the segment description text obtained after segmenting the sample video segment according to the actual segmentation points; noise injection processing is applied to the sample first description text to obtain noisy first identifier; noise injection processing is applied to the sample second description text to obtain noisy second description text; based on the sample current keyframe sequence, the noisy first description text, and the noisy second description text, the target output truth value sequence includes: a frame description truth value sequence for the sample keyframes of the sample video, a truth value sequence for the identifiers of the sample video's actual segmentation points, and a truth value sequence for the segment description text obtained after segmenting the sample video segment according to the actual segmentation points, and a truth The first noisy identifier and the noisy second descriptive text are used to determine the prediction result through the video understanding model to be trained. The prediction result includes: a predicted third descriptive text for each current keyframe in the current keyframe sequence of the sample, a predicted second identifier for the cumulative segmentation point of the cumulative keyframe sequence of the sample, and a predicted fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence of the sample according to the cumulative segmentation point. The cumulative keyframe sequence of the sample includes the current keyframe sequence of the sample and the historical keyframe sequence of the sample. Based on the predicted third descriptive text and the frame description truth sequence, the predicted second identifier and the noisy second descriptive text, the prediction result is determined by the video understanding model to be trained. The loss is calculated by identifying the ground truth sequence and the predicted fourth description text and the fragment description ground truth sequence to obtain a multi-task joint loss value; based on the predicted third description text, the predicted second identifier, and the predicted fourth description text, the sample cumulative state information is updated to obtain the sample cumulative state information of the next training time window of the current training time window; the current training time window and the next training time window partially overlap in the time dimension; based on the multi-task joint loss value, the network parameters of the video understanding model to be trained are updated to obtain the trained video understanding model; the video understanding model is used to implement the above video segmentation method.

[0025] In the above scheme, the second iterative processing module is further configured to: obtain the target frame description truth value for each current keyframe of the sample from the frame description truth value sequence; obtain the target identifier truth value of the target true segmentation point corresponding to the sample cumulative keyframe sequence from the identifier truth value sequence; obtain the target segment description truth value of each sample video segment obtained by segmenting the sample cumulative keyframe sequence according to the target true segmentation point from the segment description truth value sequence; and perform loss calculation based on the predicted third description text and the target frame description truth value, the predicted second identifier and the target identifier truth value, and the predicted fourth description text and the target segment description truth value to obtain the multi-task joint loss value.

[0026] In the above scheme, the second iterative processing module is further configured to: perform a first loss calculation based on the predicted third descriptive text and the target frame descriptive truth value to obtain a keyframe descriptive loss value; perform a second loss calculation based on the predicted second identifier and the target identifier truth value to obtain a video segmentation loss value; perform a third loss calculation based on the predicted fourth descriptive text and the target segment descriptive truth value to obtain a segment descriptive loss value; perform a fourth loss calculation based on the prediction result and the target output truth value to obtain a global loss value; and perform a weighted summation of the keyframe descriptive loss value, the video segmentation loss value, the segment descriptive loss value, and the global loss value based on preset first weights, second weights, third weights, and fourth weights to obtain the multi-task joint loss value.

[0027] In the above scheme, the second iterative processing module is further configured to: merge the one-hot encodings of the positions of each word in the target frame description ground truth in the preset dictionary into a first one-hot encoding sequence; obtain the first prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting and outputting the predicted third description text; determine the first cross-entropy loss value based on the first prediction probability in the first prediction probability distribution matrix and the first one-hot encoding label in the first one-hot encoding sequence; and determine the first cross-entropy loss value as the keyframe description loss value.

[0028] In the above scheme, the second iterative processing module is further configured to: calculate the cross-entropy loss based on the predicted second identifier and the ground truth value of the target identifier for each sample cumulative segmentation point corresponding to the sample cumulative keyframe sequence, and obtain the cross-entropy loss value; and aggregate the cross-entropy loss values ​​of all sample cumulative segmentation points corresponding to the sample cumulative keyframe sequence to obtain the video segmentation loss value.

[0029] In the above scheme, the second iterative processing module is further configured to: merge the one-hot encodings of the positions of each word in the target segment description ground truth value in the preset dictionary into a second one-hot encoding sequence; obtain the second prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting and outputting the predicted fourth description text; determine the second cross-entropy loss value based on the second prediction probability in the second prediction probability distribution matrix and the second one-hot encoding label in the second one-hot encoding sequence; and determine the second cross-entropy loss value as the segment description loss value.

[0030] In the above scheme, the second iterative processing module is further configured to: summarize the target frame description truth value, the target identifier truth value, and the target segment description truth value in the target output truth value into a target output word sequence according to a preset output format; merge the one-hot encodings of the positions of each word in the target output word sequence in the preset dictionary into a third one-hot encoding sequence; obtain the third prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting the prediction result; determine the third cross-entropy loss value based on the third prediction probability in the third prediction probability distribution matrix and the third one-hot encoding label in the third one-hot encoding sequence; and determine the third cross-entropy loss value as the global loss value.

[0031] This application provides an electronic device, including: a memory for storing computer-executable instructions or computer programs; and a processor for executing the computer-executable instructions or computer programs stored in the memory to implement the above-described video segmentation method, or to implement the above-described video understanding model training method.

[0032] This application provides a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium. When the processor of an electronic device reads the computer-executable instructions or the computer program from the computer-readable storage medium and executes the computer-executable instructions or the computer program, it implements the above-mentioned video segmentation method or the above-mentioned video understanding model training method.

[0033] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs, which are used to cause a processor to execute the computer-executable instructions or computer programs to implement the above-described video segmentation method, or to implement the above-described video understanding model training method.

[0034] The above scheme has the following beneficial effects: In this embodiment, by performing iterative processing on the keyframes of the target video according to a preset time window, and by partially overlapping the current time window with the next time window in the time dimension, the computational load of processing all keyframes in the target video at once is effectively reduced, while ensuring the continuity of information in the time dimension. During iterative processing, accumulated state information is obtained, including a first descriptive text for each historical keyframe in the historical keyframe sequence preceding the current time window, a first identifier of the historical segmentation point corresponding to the historical keyframe sequence, and a second descriptive text for each video segment obtained after segmenting the historical keyframe sequence according to the historical segmentation point. This accumulated state information is combined with the current keyframe sequence of the target video within the current time window. Based on the current keyframe sequence, the first descriptive text, the first identifier, and the second descriptive text, a semantic association between the current keyframe sequence and the historical keyframe sequence can be established, thereby accurately determining the semantic relationship between the current keyframe and the historical keyframe sequence. The sequence includes a third descriptive text for each current keyframe, a second identifier for the cumulative segmentation point corresponding to the cumulative keyframe sequence, and a fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation point. Based on the third descriptive text, the second identifier, and the fourth descriptive text, the cumulative state information is updated to obtain the cumulative state information for the next time window of the current time window. This allows the iterative processing of keyframes to dynamically adjust the cumulative state information according to the continuously acquired current keyframe sequence, thereby automatically correcting misjudgments of segmentation points or descriptive deviations caused by insufficient local keyframe information in the early stages. Finally, when completing the iterative processing of all keyframes in the target video, the target video is segmented based on the continuously updated cumulative state information obtained after the iterative processing. This effectively avoids incorrect video segmentation caused by sudden changes in local scenes or complex editing techniques, significantly improving the accuracy of the video segmentation results and the logical coherence and semantic integrity of the descriptive text corresponding to each video segment. Attached Figure Description

[0035] Figure 1 This is a diagram illustrating the video segmentation effect provided by the relevant technology; Figure 2 This is a schematic diagram of the architecture of the video segmentation system provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 4 This is another structural schematic diagram of the electronic device provided in the embodiments of this application; Figure 5 This is a flowchart illustrating the video segmentation method provided in an embodiment of this application; Figure 6 This is a schematic diagram illustrating the implementation process of obtaining keyframe sequences provided in an embodiment of this application; Figure 7 This is a schematic diagram of an optional implementation process for determining the third descriptive text, the second identifier, and the fourth descriptive text provided in an embodiment of this application; Figure 8 This is a schematic diagram of another optional implementation process for determining the third descriptive text, the second identifier, and the fourth descriptive text provided in an embodiment of this application; Figure 9 This is a schematic diagram illustrating the implementation process of determining video segmentation results provided in an embodiment of this application; Figure 10 This is a flowchart illustrating the training method of the video understanding model provided in the embodiments of this application; Figure 11 This is a schematic diagram of the implementation process for determining the joint loss value of multiple tasks provided in an embodiment of this application; Figure 12 This is a schematic diagram illustrating a fast-paced plot preview scenario provided in an embodiment of this application; Figure 13 This is another schematic diagram illustrating the plot preview scene provided in the embodiments of this application; Figure 14 This is a schematic diagram of the video plot segmentation result provided in the embodiments of this application. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0037] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this application pertain. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit the application.

[0038] If the application documents contain similar descriptions such as "first / second", the following explanation shall be added: In the following description, the terms "first / second / third" are used only to distinguish similar objects and do not represent a specific order of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0039] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0040] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0041] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0042] Related technology 1 performs scene boundary detection based on low-level visual features (such as color histograms, edge features, or motion vectors). While it can identify changes in physical scenes, it struggles to distinguish the boundaries of storylines. For example, multiple plot twists may occur within the same scene, and cross-cutting between different scenes may belong to the same plot. Because related technology 1 lacks advanced semantic understanding capabilities and scene recognition cannot replace plot segmentation, its segmentation effect is poor. For instance, the same dialogue scene may be segmented multiple times due to interludes, flashbacks, etc., and montage editing of different scenes may belong to the same plot.

[0043] Related technology 2 is a plot segmentation method based on deep learning information merging of background scenes, with video segmentation results as follows: Figure 1 As shown, four storyboards were generated: Storyboard 1, Storyboard 2, Storyboard 3, and Storyboard 4. In Storyboard 2, two frames are both road backgrounds, which are merged into the same scene. However, the first three storyboards should actually be merged together, belonging to the same scene. While related technique 2 can perform simple plot segmentation, it cannot handle varied editing techniques (such as transitions, multi-perspective plots with different scenes, and cross-cutting plots), easily producing overly fragmented segmentation effects. Therefore, it also cannot provide video segmentation effects with semantic understanding.

[0044] Related technique 3 is a method for segmenting news video programs. First, it detects and obtains feature information such as the news video's opening sequence, news title, host characteristics, camera transitions, audio silence points, switching points, and abrupt changes in pitch period. Then, based on these features, the detection results are arranged chronologically to obtain an event sequence. Next, using a predetermined symbol set and production rules, the approximate start and end positions of each news segment in the event sequence are determined. Finally, near the approximate start position, the joint posterior probability of the news segment's start position is calculated based on the event sequence. The moment with the highest posterior probability is selected as the accurate start position of the news segment, and the news video is segmented to obtain individual news video segments. Related technique 3 can effectively summarize the structural information in news videos, determine the accurate position of news segment cut points, and achieve stable and accurate segmentation. However, related technique 3 uses rule-based segmentation, which is applicable to news videos with obvious intermediate transitions (anchor speech 1 + event 1 + anchor speech 2 + event 2…), but it cannot support complex editing methods for TV dramas, movies, etc.

[0045] Related technology 4 is a video segmentation method that can improve the efficiency of video segmentation. First, the video to be segmented is acquired, which includes audio and image data. Then, the audio data in the video to be segmented is segmented based on a storyline recognition model to obtain audio segmentation results; and the image data in the video to be segmented is segmented based on a scene recognition model to obtain scene segmentation results. Next, at least one initial video segment is determined based on the audio segmentation results, and at least one scene transition frame is determined based on the scene segmentation results; finally, the video is segmented based on at least one initial video segment and at least one scene transition frame to obtain at least one target video segment. However, related technology 4, based on basic semantic understanding, cannot handle story segmentation under complex editing and varied TV / movie expression techniques: 1) It relies on audio changes, and the video segmentation effect is poor for some videos with sparse dialogue; 2) It relies on scene transition frame recognition, and the recognition effect is limited for story boundaries formed by editing methods such as one-shot, continuous shots, or direct scene transitions; 3) For situations where there are flashbacks interspersed within the same story, misidentification is prone to occur, leading to deviations in the story segmentation results.

[0046] In other words, the relevant technologies mainly suffer from the following defects: 1) The relevant technologies generally lack deep semantic understanding capabilities. The segmentation method based on visual differences can only identify changes at the picture level, such as changes in camera angle and shooting location, and it is difficult to judge the integrity of the storyline at the narrative level. Therefore, it is difficult to accurately identify the boundaries of the plot; 2) Although the audio segmentation or image content frame recognition and image segmentation frame recognition methods based on deep learning can perform plot segmentation to a certain extent, they are not adaptable to the diverse editing methods such as transitions and cross-cutting of plots from different perspectives in different scenes. They are prone to overly fragmented segmentation or incorrect segmentation positions; 3) The rule-based segmentation method is applicable to news videos with obvious structured transition content, but it is difficult to provide effective support for video content with complex narrative structures and complex editing methods, such as TV series and movies.

[0047] This application addresses at least one problem in related technologies by proposing a video segmentation method that constructs a processing flow combining scene segmentation and large-scale model plot understanding. In the scene segmentation stage, computer vision technology combined with a lightweight model is used to analyze the video scene segmentation. Video frames are extracted using a fine-grained sampling rate to cover rapidly changing action and dialogue information. Furthermore, a scene classification model and an Optical Character Recognition (OCR) model are used to extract scene embedding features and text from the video frames in parallel, and scene segmentation is performed based on the similarity of the scene embedding features. On this basis, redundant frames with similar visuals but repeated dialogue are removed through text comparison to obtain the scene segmentation results and representative keyframes for each scene. This processing method can improve the representativeness and stability of keyframe selection while controlling computational complexity. In the plot understanding stage, a cumulative understanding and correction mechanism is used to progressively analyze the keyframe sequence. Specifically, keyframe sequences are input into a large model, and combined with reference character information, the model describes the plot of the keyframe content, merges plot relationships, and adjusts the cumulative segmentation positions. As keyframe sequences are gradually input, the large model continuously updates its understanding of the context based on cumulative forward inference, and corrects previously formed plot segmentation and description results based on newly acquired information, thereby gradually obtaining more accurate plot boundaries and a more complete plot summary. This video segmentation method can not only identify explicit scene transitions, but also handle plot twists within the same scene, continuation of the same plot across scenes, and complex expressive forms such as interludes, flashbacks, and cross-narratives.

[0048] Here, we first describe an exemplary application of the video segmentation device in this application embodiment. This video segmentation device is an electronic device used to implement a video segmentation method. The video segmentation device (i.e., electronic device) provided in this application embodiment can be implemented as a terminal or as a server. In one implementation, the electronic device provided in this application embodiment can be implemented as any terminal with video segmentation function or video understanding model training function, such as a laptop, tablet, desktop computer, or intelligent robot. In another implementation, the electronic device provided in this application embodiment can also be implemented as a server, wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in this application embodiment. Below, we will describe an exemplary application when the electronic device is implemented as a server.

[0049] See Figure 2 , Figure 2 This is a schematic diagram of the structure of the video segmentation system 100 provided in the embodiment of this application. In order to support a video segmentation application, the terminal 400 connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0050] Terminal 400 is used to send a video segmentation request for the target video to server 200. Server 200 constitutes the video segmentation device of this application embodiment. In response to the video segmentation request, server 200 obtains the video segmentation result of the target video through the video segmentation method provided in this application embodiment, and returns the video segmentation result of the target video to terminal 400 so as to realize the output of the video segmentation result of the target video on terminal 400.

[0051] In some embodiments, server 200 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0052] In some embodiments, a training system for implementing a training method for a video understanding model is also provided. This training system may be the same system as the video segmentation system described above, or it may be a different system. In this training system, a terminal sends a training request for the video understanding model to a server. The server constitutes the training device for the video understanding model in this embodiment of the application. In response to the training request, the server trains the video understanding model using the training method provided in this embodiment of the application, obtains the model parameters of the trained video understanding model, and returns the model parameters to the terminal, so as to enable the terminal to call the trained video understanding model to perform video segmentation on the target video.

[0053] See Figure 3 , Figure 3 This is a schematic diagram of the structure of the electronic device 40 provided in the embodiment of this application. Figure 3 The illustrated electronic device 40 can be a video segmentation device, or it can be a training device for a video understanding model. The electronic device includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 3 The general labeled all buses as Bus System 440.

[0054] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0055] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0056] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0057] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0058] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0059] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, WiFi, and Universal Serial Bus (USB); the presentation module 453 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.); the input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.

[0060] In some embodiments, the apparatus provided in this application may be implemented in software. Figure 3 A video segmentation device 455 stored in memory 450 is shown. The video segmentation device 455 can be a video segmentation device in an electronic device, and can be software in the form of programs and plug-ins, including the following software modules: a first iterative processing module 4551 and a video segmentation module 4552. These modules are logically related and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.

[0061] In other embodiments, the apparatus provided in this application can be implemented in software. Figure 4A training device 456 for a video understanding model stored in memory 450 is shown. The training device 456 can be a training device for a video understanding model in an electronic device, and can be software in the form of programs and plugins, including the following software module: a second iterative processing module 4561. These modules are logically related and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.

[0062] In some embodiments, the apparatus provided in this application can be implemented in hardware. For example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the video segmentation method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0063] The video segmentation methods provided in the embodiments of this application can be executed by an electronic device, which can be a terminal or a server. That is, the video segmentation methods in the embodiments of this application can be executed by a terminal, by a server, or by interaction between a terminal and a server.

[0064] The video segmentation method provided in the embodiments of this application will be described below. Figure 5 This is a flowchart illustrating the video segmentation method provided in the embodiments of this application. The following will be combined with... Figure 5 The steps shown are explained as follows: Figure 5 As shown, taking the server as the execution subject of the video segmentation method as an example, this video segmentation method performs the following iterative processing on the key frames of the target video according to a preset time window through steps S101 to S104: Step S101: Obtain the current keyframe sequence and cumulative status information of the target video within the current time window; the cumulative status information includes: the first description text of each historical keyframe in the historical keyframe sequence before the current time window, the first identifier of the historical segmentation point corresponding to the historical keyframe sequence, and the second description text of each video segment obtained after segmenting the historical keyframe sequence according to the historical segmentation point.

[0065] In this embodiment, the target video is the video data object for which video segmentation needs to be performed. The target video can be an episode of a TV series, a movie, or a segment of online video content. The target video is the source of all keyframes and also the carrier of the final video segmentation result. For example, a 45-minute episode file can be used as the target video. Keyframes are frames selected from the target video that represent changes in video content near a given time position; keyframes serve as the foundational data for subsequent analysis. The server processes keyframes, rather than all frames of the target video frame by frame, which reduces the amount of data while preserving the main content changes.

[0066] A preset time window is a time range division rule determined before the start of iterative processing. The preset time window is used to limit the range of keyframes covered by each round of iterative processing. The preset time window can be represented by a start time and an end time, or by the time interval where the keyframes are located. The purpose of the preset time window is to divide the target video into several continuously processable local ranges, keeping the data scale faced by the server controllable in each round of processing, while preserving the temporal sequence. Furthermore, every two adjacent time windows partially overlap in the time dimension, meaning that every two adjacent time windows have at least one overlapping keyframe. Iterative processing is a method of repeatedly executing processing actions in chronological order. The intermediate results formed in the previous round of iterative processing are incorporated into the next round of iterative processing, and the next round of iterative processing continues based on the previous round. The technical significance of iterative processing is that the server does not need to process all keyframes at once, but rather processes them in rounds and gradually accumulates the results. For example, the first time window is processed first, then the second time window, and so on, until all time windows are processed.

[0067] For example, if the target video is a 120-second narrative video, the server extracts 20 keyframes from the target video, numbered 1 to 20 in chronological order, with each keyframe spaced approximately 0.2 seconds apart. The content of the target video is as follows: Keyframes 1 to 4 show a panoramic view of the city, high-rise buildings, roads, and overpasses, setting the scene before the characters arrive in the city; Keyframes 5 to 8 show a woman entering the woods and observing her surroundings; Keyframes 9 to 12 show the woman stopping in the woods and speaking, with the caption "Sometime three years ago" appearing; Keyframes 13 to 16 show the scene transitioning to a classroom, where a man and woman are talking, depicting a flashback; Keyframes 17 to 20 show the scene returning to the woods, where the woman continues her journey and discovers clues. The server sets up a time window of 8 keyframes, with two adjacent time windows overlapping by 2 keyframes. Therefore, the preset time windows are: Keyframes 1 to 8, Keyframes 7 to 14, and Keyframes 13 to 20.

[0068] In some embodiments, the server reads keyframes from the target video in chronological order and defines the first processing range according to a preset time window, obtaining the first iteration processing range, i.e., the time window for the first iteration processing. Then, within the time window of the first iteration processing, keyframes are retrieved, and the iteration processing flow is started to obtain the result of the first iteration processing. Next, based on the result of the first iteration processing, the preset time window is moved chronologically, and the same iteration processing flow is repeated within the new processing range to obtain the results of subsequent iteration processing. Finally, it is checked whether the results of subsequent iteration processing have covered all keyframes in the target video, and the iteration processing is completed when all keyframes are covered. For example, for the same target video, the server first divides the keyframe processing range according to a preset time window. Keyframes 1 to 8 are determined as the first time window, keyframes 7 to 14 are determined as the second time window, and keyframes 13 to 20 are determined as the third time window. Therefore, when performing iterative processing on the keyframes of the target video, keyframes 1 to 20 are not processed all at once. Instead, keyframes 1 to 8 are processed first, followed by keyframes 7 to 14, and finally keyframes 13 to 20. Since keyframes 7 to 8 are located in both the first and second time windows, and keyframes 13 to 14 are located in both the second and third time windows, the server can repeatedly observe the keyframe content near the boundary at the time window intersection.

[0069] In this embodiment, the current time window is the time window actually selected in the current iteration. The current time window defines the range of keyframes directly processed in this round. The current keyframe sequence is a set of keyframes located within the current time window and arranged chronologically; the current keyframe sequence is the direct object of this iteration. The cumulative state information is a set of historical processing results retained before entering the current time window. The cumulative state information is not a single piece of data, but a data set containing multiple levels of information, used to pass the analysis results from previous time ranges to the current time window. The cumulative state information includes: a first descriptive text for each historical keyframe in the historical keyframe sequence before the current time window, a first identifier of the historical segmentation point corresponding to the historical keyframe sequence, and a second descriptive text for each video segment obtained after segmenting the historical keyframe sequence according to the historical segmentation point. The historical keyframe sequence is a set of keyframes located before the current time window and arranged chronologically. The historical keyframe sequence reflects all keyframes that have entered the processing range before the current time window. A historical keyframe is any keyframe in the historical keyframe sequence, and each historical keyframe can have an independent text description. The first descriptive text is the text description already formed for each historical keyframe. Its purpose is to transform the content information of the historical keyframes into text information that can be used in subsequent steps. A historical breakpoint is a determined breakpoint in the historical keyframe sequence. There can be one, multiple, or no breakpoints. When a historical breakpoint is empty, it indicates that the breakpoint has not yet been determined within the historical keyframe sequence. For example, the current time window is keyframes 7 to 14, and the historical keyframe sequence is keyframes 1 to 8. If the breakpoint has been determined at keyframe 5, then the historical breakpoint is keyframe 5; if keyframes 1 to 8 are determined to belong to the same video scene, then the historical breakpoint is empty. The first identifier is a position marker obtained after recording the breakpoint position of the historical breakpoint. The first identifier can be represented by a keyframe number, time position, or sequence number, and it uniquely indicates the position of the historical breakpoint. When a historical breakpoint is empty, the first identifier is also empty. For example, when the historical segmentation point is keyframe 5, the first identifier can be recorded as [5]; when the historical segmentation point is empty, the first identifier can be recorded as an empty list. A video segment is a continuous time period formed by segmenting the keyframe sequence according to the segmentation point. For example, when the historical segmentation point is keyframe 5, keyframes 1 to 4 can form the first video segment, and keyframes 5 to 8 can form the second video segment. The second descriptive text is a text description that has been formed for each video segment, and the second descriptive text reflects the content summary at the video segment level.

[0070] In some embodiments, the server reads keyframes within the current time window from the target video and arranges them chronologically to obtain the current keyframe sequence. Then, it extracts accumulated state information from historical processing records saved prior to the current time window to obtain accumulated state information used in this iteration. Next, it extracts the description record of each historical keyframe from the accumulated state information to obtain a first description text; it extracts the segmentation position records from the accumulated state information and marks them to obtain a first identifier; it extracts the description record of each video segment formed after segmenting the historical keyframe sequence according to historical segmentation points from the accumulated state information to obtain a second description text. Finally, it summarizes the current keyframe sequence, the first description text, the first identifier, and the second description text to obtain the dataset used in this iteration.

[0071] In some embodiments, when the current time window is the first time window within a preset time window, initialization can be completed with empty historical processing records, resulting in empty accumulated state information, i.e., the first description text, the first identifier, and the second description text are all empty. For example, when the current time window is the first time window, i.e., keyframes 1 to 8, since there are no earlier keyframes before keyframes 1 to 8, the historical keyframe sequence before the current time window is empty, and therefore the accumulated state information is also empty. At this time, the obtained current keyframe sequence is keyframes 1 to 8, and the first description text, the first identifier, and the second description text in the obtained accumulated state information are all empty. When the current time window is the second time window, i.e., keyframes 7 to 14, the obtained current keyframe sequence is keyframes 7 to 14. Since keyframes 1 to 8 are located before the current time window, keyframes 1 to 8 are used as the historical keyframe sequence.

[0072] After the previous round of processing, the first description text has been formed for key frames 1 to 8. For example: the first description text of key frame 1 is "showing the exterior view of the city's high-rise buildings"; the first description text of key frame 2 is "showing the busy roads and traffic flow"; the first description text of key frame 3 is "showing the overhead view of the overpass"; the first description text of key frame 4 is "showing the panoramic view of the city"; the first description text of key frame 5 is "the woman enters the woods"; the first description text of key frame 6 is "the woman stops and observes in the woods"; the first description text of key frame 7 is "the woman continues to walk in the woods"; the first description text of key frame 8 is "the woman stops and observes ahead". If it is determined that key frames 1 to 4 and key frames 5 to 8 belong to two video segments, then the historical segmentation point is key frame 5, and the first identifier is [5]. After segmenting key frames 1 to 8 according to key frame 5, two video segments are obtained. The second descriptive text for the video clips with keyframes 1 to 4 is “Urban environment setup”, and the second descriptive text for the video clips with keyframes 5 to 8 is “Women enter the woods and gradually unfold their actions”.

[0073] When the current time window is the third time window, i.e., keyframes 13 to 20, the acquired current keyframe sequence is keyframes 13 to 20. The historical keyframe sequence preceding the current time window is keyframes 1 to 14. The server reads the first description text of each of keyframes 1 to 14 from the data retained in the previous two rounds, reads the first identifier of the historical segmentation point, and then reads the second description text of each video segment obtained after segmentation according to the historical segmentation point, as the accumulated state information used in this round of processing.

[0074] In some embodiments, see Figure 6 , Figure 6 This demonstrates that before acquiring the current keyframe sequence and cumulative state information of the target video within the current time window, the following steps S201 to S204 can also be performed: Step S201: Extract frames from the target video according to the preset sampling rate to obtain an initial video frame sequence.

[0075] The preset sampling rate refers to the frequency of video frame extraction set before frame extraction. The preset sampling rate limits how many video frames are extracted from the target video per unit of time. The preset sampling rate can be expressed as the number of video frames extracted per second. For example, a preset sampling rate of 5 frames / second means 5 video frames are extracted from the target video per second. An initial video frame refers to a single video frame extracted from the target video according to the preset sampling rate. An initial video frame is a discrete image representation of the target video at a given time position; each initial video frame can be associated with a specific time position. For example, a 120-second target video at a sampling rate of 5 frames / second can yield multiple initial video frames, and an initial video frame at the 10th second can represent the video content around the 10th second. An initial video frame sequence refers to an ordered set formed by arranging all initial video frames in chronological order. The initial video frame sequence reflects the temporal structure of the target video after sampling and is the foundational data for subsequent processing. Each position in the initial video frame sequence corresponds to an initial video frame. For example, a 120-second target video can be extracted at 5 frames per second to obtain 600 initial video frames, which can be arranged in chronological order to form an initial video frame sequence.

[0076] In some embodiments, the target video and a preset sampling rate are used as input, video frames are selected in chronological order along the target video according to the preset sampling rate, the selection order is recorded, and an initial video frame sequence arranged in chronological order is output.

[0077] Step S202: Determine the scene embedding features and text information of each initial video frame in the initial video frame sequence.

[0078] Scene embedding features refer to feature data used to characterize the content of the initial video frame. Scene embedding features can reflect differences in scene environment, character distribution, composition, or content within the initial video frame. Scene embedding features can be represented in vector form to facilitate comparison and differentiation between different initial video frames. For example, the scene embedding features of a city road scene may be significantly different from those of a forest scene; however, the scene embedding features of consecutive shots within the same forest may be quite similar. Text information refers to the text content contained within the initial video frame. Text information can be subtitles or other text content within the initial video frame. Text information supplements the visual information with additional textual clues, allowing the server to utilize both image and textual information simultaneously when analyzing the content of the initial video frame.

[0079] In some embodiments, an initial video frame sequence is taken as input, scene embedding features representing the content of the scene are extracted for each initial video frame in the initial video frame sequence, and text information in the initial video frame is extracted, and scene embedding features and text information of each initial video frame are output.

[0080] Step S203: Based on scene embedding features, the initial video frame sequence is segmented into multiple video segments.

[0081] A video storyboard refers to a continuous set of initial video frames obtained by segmenting an initial video frame sequence. The initial video frames in a video storyboard are sequential in time and have a high degree of consistency in content. Multiple video storyboards together cover the entire time range of the initial video frame sequence. For example, a set of continuous initial video frames showing an urban environment can form a video storyboard; a set of continuous initial video frames showing action in a forest can form another video storyboard.

[0082] In some embodiments, the initial video frame sequence is segmented into scenes based on scene embedding features, which can be achieved as follows: First, the feature similarity between the scene embedding features of each of the two adjacent initial video frames in the initial video frame sequence is determined; then, the feature similarity is normalized to obtain a normalized feature similarity; next, when the normalized feature similarity is less than a preset first feature similarity threshold, the later initial video frame among the two adjacent initial video frames is determined as the scene segmentation point; finally, the initial video frame sequence is segmented into scenes based on the scene segmentation point to obtain multiple video scenes.

[0083] Feature similarity refers to the numerical value of the similarity between the scene embedding features of two adjacent initial video frames. Feature similarity is used to indicate how similar the scene embedding features of two initial video frames are in terms of content. A higher feature similarity value indicates that the scene embedding features of the two initial video frames are more similar; a lower feature similarity value indicates that the scene embedding features of the two initial video frames are more different.

[0084] In some embodiments, the initial video frame sequence and the scene embedding features of each initial video frame in the initial video frame sequence are taken as input, and each pair of adjacent initial video frames in the initial video frame sequence is traversed in chronological order. The proximity between the two scene embedding features is calculated, and the feature similarity between the scene embedding features of each pair of adjacent initial video frames is output.

[0085] In some embodiments, the feature similarity is normalized to obtain the normalized feature similarity, which can be achieved as follows: First, a reference video frame sequence of a preset duration is extracted from the target video; second, the reference feature similarity between the scene embedding features of each two adjacent reference video frames in the reference video frame sequence is determined; then, the mean similarity and standard deviation of the multiple reference feature similarities are determined; next, for the feature similarity of each two adjacent initial video frames in the initial video frame sequence, the difference between the feature similarity and the mean similarity is determined; finally, the ratio of the difference to the standard deviation of the similarity is determined as the normalized feature similarity.

[0086] Here, extraction refers to the process of selecting video frames from the target video that fall within a preset duration range and forming an ordered result. Extraction is not random reading, but rather selection is performed according to a time range and time order. The extracted result directly serves as the data source for subsequent processing. For example, video frames within the first 24 seconds of a 120-second target video can be selected to form a reference video frame sequence; alternatively, video frames from the 10th to the 34th second of the target video can be selected to form a reference video frame sequence. The preset duration refers to the time length set before the extraction action is performed. The preset duration is used to limit the time range covered by the reference video frame sequence. The preset duration can be expressed in seconds, minutes, or other time units. The purpose of setting the preset duration is to allow the server to first obtain a set of video frames within a fixed time range from the target video, and then form the data basis for subsequent calculations based on this set of video frames within a fixed time range. The reference video frame sequence refers to the set of video frames extracted from the target video and arranged in chronological order by the server. The reference video frame sequence is used to reflect the changes in the target video within the preset duration range. The difference between the reference video frame sequence and the initial video frame sequence is that the reference video frame sequence is specifically used to form the statistical basis required for subsequent normalization processing, while the initial video frame sequence is used to cover the overall processing range of the target video.

[0087] In some embodiments, a target video and a preset duration are used as input. Video frames within the preset duration range are selected from the target video in chronological order, and the selected video frames are sorted by time to output a reference video frame sequence.

[0088] Reference feature similarity refers to the numerical value of the similarity between the scene embedding features of each of two adjacent reference video frames. Reference feature similarity is used to characterize the proximity between adjacent frames within a sequence of reference video frames. For example, the reference feature similarity between two consecutive city frames can be 0.95, and the reference feature similarity between a city frame and a forest frame can be 0.41.

[0089] In some embodiments, the reference video frame sequence and the scene embedding features of each reference video frame in the reference video frame sequence are taken as input, and every two adjacent reference video frames in the reference video frame sequence are traversed in chronological order. The similarity between the two scene embedding features is calculated, and multiple reference feature similarities are output.

[0090] The mean similarity is the average value obtained after statistically analyzing the similarities of multiple reference features. It characterizes the overall level of closeness between adjacent frames within a reference video frame sequence. A higher mean similarity indicates that adjacent frames in the reference video frame sequence are generally closer; a lower mean similarity indicates that adjacent frames in the reference video frame sequence exhibit more significant variations. The standard deviation of similarity is the fluctuation value obtained after discretely analyzing the similarities of multiple reference features. It characterizes the degree of dispersion of the distribution of the similarities of multiple reference features around the mean similarity. A larger standard deviation indicates greater fluctuation in the variation amplitude of adjacent frames in the reference video frame sequence; a smaller standard deviation indicates a more concentrated variation amplitude among adjacent frames in the reference video frame sequence.

[0091] In some embodiments, multiple reference feature similarities are taken as input, and centralized and discrete statistics are performed on the multiple reference feature similarities to output the mean similarity and the standard deviation of similarity.

[0092] The difference is the numerical result obtained by subtracting the mean similarity from the similarity of a given feature. The difference indicates the degree of deviation of a single feature's similarity from the reference statistical center. A positive difference indicates that the feature's similarity is higher than the mean similarity; a negative difference indicates that the feature's similarity is lower than the mean similarity; the larger the absolute value of the difference, the greater the degree of deviation.

[0093] In some embodiments, the feature similarity and mean similarity of each pair of adjacent initial video frames in the initial video frame sequence are used as input, and the deviation value between the feature similarity and the mean similarity is calculated for each feature similarity, and multiple differences are output.

[0094] The ratio is the numerical result obtained by dividing the difference by the standard deviation of similarity. The ratio is used to convert the deviation of individual feature similarity to a uniform scale. A larger ratio indicates a more significant degree that the similarity of a single feature is higher than the mean similarity; a smaller ratio indicates a more significant degree that the similarity of a single feature is lower than the mean similarity.

[0095] In some embodiments, multiple differences and similarity standard deviations are used as inputs. Each difference is divided by the similarity standard deviation, and the resulting ratio is determined as the normalized feature similarity. Multiple normalized feature similarities are then output.

[0096] Here, a reference video frame sequence of a preset duration is first extracted from the target video. Then, the reference feature similarity between adjacent reference video frames within the reference video frame sequence is determined, and the mean similarity and standard deviation of the similarity are further obtained. Subsequently, the difference between the feature similarity in the initial video frame sequence and the mean similarity is calculated. Finally, the ratio of the difference to the standard deviation of the similarity is determined as the normalized feature similarity. This allows the normalized feature similarity to be directly based on the reference statistical data of the target video itself, which can reduce the impact of differences in picture style, similarity distribution, and local numerical scale between different target videos on subsequent judgments. This improves the comparability, consistency, and stability of the normalized feature similarity within the same target video and between different target videos.

[0097] In some embodiments, the preset first feature similarity threshold refers to a comparison benchmark value set before the judgment is performed. The preset first feature similarity threshold is used to compare with the normalized feature similarity to determine whether there is a change between two adjacent initial video frames sufficient to trigger a scene segmentation. When the normalized feature similarity is lower than the preset first feature similarity threshold, it indicates that the scene change between two adjacent initial video frames meets the preset judgment condition. The scene segmentation point refers to the position of the initial video frame in the initial video frame sequence that is identified as the starting position of the video segment. The scene segmentation point is used to mark the boundary between video scenes. The later initial video frame among two adjacent initial video frames is used as the scene segmentation point, so the scene segmentation point directly points to the starting position of the new video scene.

[0098] In some embodiments, the normalized feature similarity and the preset first feature similarity threshold are used as inputs. The relationship between each normalized feature similarity and the preset first feature similarity threshold is compared one by one. When the normalized feature similarity is less than the preset first feature similarity threshold, the later initial video frame in the two adjacent initial video frames corresponding to the normalized feature similarity is determined as the segmentation point, and one or more segmentation points are output.

[0099] In some embodiments, the initial video frame sequence and the segmentation points are taken as input, the segmentation points are inserted into the corresponding positions of the initial video frame sequence in chronological order, and the initial video frame sequence is segmented into segments using the segmentation points as boundaries to output multiple video segments.

[0100] Here, by first determining the feature similarity between the scene embedding features of each of the two adjacent initial video frames, then normalizing the feature similarity, and when the normalized feature similarity is less than a preset first feature similarity threshold, the later initial video frame among the two adjacent initial video frames is determined as the segmentation point. Finally, the segmentation is completed based on the segmentation point. This makes the determination of the segmentation boundary based on the common basis of changes in adjacent scenes and a unified numerical scale, which can reduce the impact of feature similarity value fluctuations caused by different video content and different time positions on the segmentation results. At the same time, it makes the positioning rules of the segmentation point clear, the segmentation boundary clear, and the segmentation results highly consistent, thereby improving the interpretability and processing stability of multiple video segments.

[0101] Step S204: Based on scene embedding features and text information, deduplication is performed on the initial video frames in each video segment to obtain the keyframe sequence of the target video.

[0102] Deduplication refers to the process of identifying initial video frames with duplicate content or information within each video shot, based on scene embedding features and text information, and retaining representative initial video frames from these duplicates. The purpose of deduplication is to reduce the number of redundant initial video frames and retain those that represent changes in the video shot's content. For example, in the same video shot, if multiple consecutive initial video frames depict a woman standing in a forest with identical text information, only the representative initial video frames can be retained; similarly, if two initial video frames in the same video shot have similar visuals but different text information, each initial video frame can be retained separately.

[0103] In some embodiments, based on scene embedding features and text information, the initial video frames in each video segment are deduplicated to obtain the keyframe sequence of the target video. This can be achieved as follows: First, for each video segment, when any two adjacent initial video frames in the video segment meet at least one of the following preset conditions, the text quality of the text information of each of the two adjacent initial video frames is determined; wherein, the preset conditions are: the feature similarity between the scene embedding features of each of the two adjacent initial video frames in the video segment is greater than a preset second feature similarity threshold; the text similarity between the text information of each of the two adjacent initial video frames is greater than a preset text similarity threshold; then, when the text quality corresponding to the earlier initial video frame in any two adjacent initial video frames is greater than or equal to a preset quality threshold, the earlier initial video frame is retained; when the text quality corresponding to the earlier initial video frame is less than the quality threshold, the initial video frame corresponding to the highest text quality in any two adjacent initial video frames is retained; finally, the retained initial video frames are sorted according to time order to obtain the keyframe sequence of the target video.

[0104] Preset conditions refer to the set of judgment conditions set before performing the text quality determination action. Preset conditions are used to filter any two adjacent initial video frames within a video storyboard that need to enter the deduplication judgment. The preset conditions include two conditions: first, the feature similarity between the scene embedding features of any two adjacent initial video frames is greater than a preset second feature similarity threshold; second, the text similarity between the text information of any two adjacent initial video frames is greater than a preset text similarity threshold. As long as either of these two conditions is met, the text quality of the text information of any two adjacent initial video frames is determined. For example, if two initial video frames have essentially the same visuals but different text information, as long as the feature similarity is greater than the preset second feature similarity threshold, the text quality determination process begins; similarly, if two initial video frames have slight visual differences but highly consistent text information, as long as the text similarity is greater than the preset text similarity threshold, the text quality determination process also begins. The preset second feature similarity threshold is a pre-set numerical benchmark used to compare feature similarity. The preset second feature similarity threshold is used to determine whether any two adjacent initial video frames are sufficiently similar in content. When the feature similarity is greater than the preset second feature similarity threshold, it indicates that the two initial video frames have reached the level of visual content suitable for deduplication analysis. Text similarity refers to the degree of closeness between the text information of any two adjacent initial video frames. Text similarity is used to measure whether the text information in two initial video frames expresses the same or similar text content. The preset text similarity threshold is a pre-set numerical benchmark for comparing text similarity. The preset text similarity threshold is used to determine whether the text information in any two adjacent initial video frames is sufficiently close. When the text similarity is greater than the preset text similarity threshold, it indicates that the text information in the two initial video frames has reached the level of suitable for deduplication analysis. Text quality refers to the usability of text information. Text quality characterizes the retention value of text information in deduplication processing. Higher text quality indicates more complete, clearer, or more suitable text information for subsequent keyframe retention; lower text quality indicates more missing text information, less effective content, or weaker usability.

[0105] In some embodiments, the initial video frame in each video segment, the scene embedding features of each initial video frame, and the text information of each initial video frame are used as inputs. The process checks whether any two adjacent initial video frames in each video segment meet at least one preset condition. When at least one preset condition is met, the text quality of the text information of each of the two initial video frames is determined. The text quality of any two adjacent initial video frames that meet at least one preset condition is output.

[0106] The preceding initial video frame refers to the initial video frame that appears earlier in time among any two adjacent initial video frames. The preceding initial video frame is used as the priority selection object. By first checking the text quality of the preceding initial video frame, a priority retention rule is formed for adjacent duplicate initial video frames. The preset quality threshold is a pre-set numerical benchmark used to compare text quality. The preset quality threshold is used to determine whether the text quality of the preceding initial video frame meets the requirements for direct retention. When the text quality is greater than or equal to the preset quality threshold, it indicates that the preceding initial video frame has sufficient retention value.

[0107] In some embodiments, the text quality of any two adjacent initial video frames and a preset quality threshold are used as input. For each pair of any two adjacent initial video frames, the text quality of the preceding initial video frame is checked to see if it is greater than or equal to the preset quality threshold. If the check result is yes, the preceding initial video frame is retained, and the retention result is output.

[0108] The initial video frame with the highest text quality refers to the one with the higher text quality value among any two adjacent initial video frames. The initial video frame with the highest text quality is used as a substitute for the preceding initial video frame when it cannot be directly retained. The initial video frame with the highest text quality can be either the preceding or following initial video frame.

[0109] In some embodiments, the text quality of any two adjacent initial video frames and a preset quality threshold are used as input. For the preceding initial video frame whose text quality is less than the preset quality threshold, the text quality of any two adjacent initial video frames is compared, and the initial video frame with the highest text quality is retained. The retained result is output. The retained initial video frames have eliminated redundant initial video frames in the video storyboard, therefore, the retained initial video frames are the direct data source for forming the keyframe sequence.

[0110] In some embodiments, the initial video frames retained within each video segment can be used as input, and the retained initial video frames can be sorted according to their temporal order in the target video to output the keyframe sequence of the target video.

[0111] Here, within each video segment, preset conditions for entry determination are first set based on feature similarity and text similarity. Then, the text quality of any two adjacent initial video frames that meet at least one preset condition is determined. Priority is given to retaining the earlier initial video frame if its text quality reaches a preset quality threshold. If the text quality of the earlier initial video frame does not reach the preset quality threshold, the initial video frame with the highest text quality is retained. Finally, the keyframe sequence of the target video is formed in chronological order. This allows the deduplication process to consider the similarity of images, text similarity, usability of text information, and chronological order, thereby reducing redundant initial video frames within the video segment, maintaining the quality and temporal consistency of text information in the keyframe sequence, and improving the representativeness, stability, and applicability of the keyframe sequence of the target video to the video content.

[0112] Here, by extracting frames from the target video according to a preset sampling rate, the continuous video data is first converted into an initial video frame sequence. Then, the scene embedding features and text information of each initial video frame are determined, and multiple video sub-scenes are formed based on the scene embedding features. Subsequently, deduplication is performed within each video sub-scene based on the scene embedding features and text information. This enables the keyframe sequence obtained by the server to simultaneously achieve the technical effects of clear temporal order, clear image representativeness, complete preservation of text clues, and controlled number of redundant frames. This reduces the data scale pressure caused by directly processing the target video frame by frame, improves the ability of the keyframe sequence to represent changes in the content of the target video, and enhances the consistency, processability, and result stability of the keyframe sequence.

[0113] Step S102: Based on the current keyframe sequence, the first description text, the first identifier, and the second description text, determine the third description text for each current keyframe in the current keyframe sequence, the second identifier of the cumulative segmentation point corresponding to the cumulative keyframe sequence, and the fourth description text for each video segment obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation point; the cumulative keyframe sequence includes the current keyframe sequence and the historical keyframe sequence.

[0114] The third descriptive text is a text description generated for each current keyframe in the current keyframe sequence. The third descriptive text is the same as the first descriptive text at the data level, both being keyframe-level text information, but the third descriptive text is generated for the current keyframe sequence. The cumulative keyframe sequence is a keyframe sequence formed by connecting the historical keyframe sequence and the current keyframe sequence in chronological order. The cumulative split point is the splitting position determined in the cumulative keyframe sequence. The cumulative split point can not only retain historical split points, but can also adjust the splitting positions near the historical time range or add new splitting positions based on changes in the content of the current keyframe sequence. The second identifier is a position marker obtained by recording the splitting positions of the cumulative split points. The second identifier can be the same as the first identifier in terms of representation, but the second identifier is oriented towards the cumulative keyframe sequence. For example, the second identifier can be recorded as [5, 9, 16]. The fourth descriptive text is a text description generated for each video segment formed after splitting according to the cumulative split points. The fourth descriptive text is oriented towards the segment-level results on the cumulative keyframe sequence.

[0115] For example, when the current time window is the first time window, the server determines the third description text of each current key frame based on the current key frame sequence from key frame 1 to key frame 8. For example: the third description text of key frame 1 is "city high-rise buildings appear in a sunny environment"; the third description text of key frame 2 is "road vehicles pass quickly"; the third description text of key frame 3 is "the overpass structure scene cuts in"; the third description text of key frame 4 is "the city panorama continues to be displayed"; the third description text of key frame 5 is "the woman enters the forest"; the third description text of key frame 6 is "the woman turns to observe the surroundings"; the third description text of key frame 7 is "the woman continues to walk in the forest"; the third description text of key frame 8 is "the woman stops and prepares to speak". Key frames 1 to key frames 8 are taken as the cumulative key frame sequence, and after analysis, key frame 5 is determined to be the cumulative cut point, so the second identifier is [5]. Next, keyframes 1 to 8 are divided according to keyframe 5 to form two video segments: keyframes 1 to 4 form the first video segment, with the fourth descriptive text being "continuous display of urban exterior scenery, serving as environmental setup"; keyframes 5 to 8 form the second video segment, with the fourth descriptive text being "the woman enters the woods and gradually unfolds her actions".

[0116] When the current time window is the second time window, after the server enters the current time window of keyframes 7 to 14, it continues to determine the third description text of keyframes 7 to 14 based on the current keyframe sequence, the first description text, the first identifier, and the second description text. For example: the third description text of keyframe 7 is "The woman continues walking in the woods"; the third description text of keyframe 8 is "The woman stops and looks ahead"; the third description text of keyframe 9 is "The woman speaks, and the subtitle shows some time three years ago"; the third description text of keyframe 10 is "The woman is still talking in the woods"; the third description text of keyframe 11 is "The scene begins to hint at a memory transition"; the third description text of keyframe 12 is "The memory transition becomes clearer"; the third description text of keyframe 13 is "The scene enters the classroom, and the man and woman look at each other"; the third description text of keyframe 14 is "The classroom dialogue begins to unfold". The server merges the historical keyframe sequence of keyframes 1 to 8 with the current keyframe sequence of keyframes 7 to 14 into a cumulative keyframe sequence: keyframes 1 to 14. After re-evaluating the cumulative segmentation points on the cumulative keyframe sequence, the server found that keyframe 5 should still be retained, and added a new cumulative segmentation point at keyframe 13, because keyframe 13 marks the transition from the real scene in the forest to the memory scene in the classroom. Therefore, the second identifier is updated to [5,13]. The server then segments keyframes 1 to 14 according to keyframes 5 and 13, resulting in three video segments: keyframes 1 to 4 form the first video segment, with the fourth descriptive text being "city environment setup"; keyframes 5 to 12 form the second video segment, with the fourth descriptive text being "the woman moves in the forest and introduces a memory through dialogue"; and keyframes 13 to 14 form the third video segment, with the fourth descriptive text being "the scene enters the classroom, and the memory plot begins."

[0117] In some embodiments, see Figure 7 , Figure 7 This shows that step S102 can be achieved by performing the following steps S1021 to S1024: Step S1021: Obtain the object reference information corresponding to the target video, and the text information corresponding to each current keyframe in the current keyframe sequence.

[0118] The objects in the target video refer to the main subjects in the frame that need to be identified when processing the current keyframe sequence. These objects can be characters or other objects that contribute to understanding the plot. Objects can include characters (e.g., female character A or male character B) and non-character objects that indicate the plot (e.g., cars, cell phones, or identification documents). Object reference information refers to the set of reference information that establishes a connection with objects in the target video and is used for semantic recognition within the current time window. Object reference information provides a stable basis for object recognition. Object reference information can include at least one of the following: object name, object image, description of the relationship between objects, and description of object appearance features. The role of object reference information is to first establish object recognition references when processing the current keyframe sequence, and then combine this with the content of the frame in the current keyframe sequence for understanding, thereby reducing object confusion. For example, in a 120-second target video, object reference information includes "Object A: Frontal image of a female object, object name: Female Object A," and "Object B: Frontal image of a male object, object name: Male Object B." Text information refers to the text content that establishes a correspondence with each current keyframe. Text information is used to supplement explicit semantic cues in the current keyframe. Text information can include at least one of the following: subtitles, dialogue, or visual prompts. The purpose of text information is to provide textual evidence beyond the image for the current keyframe, allowing the server to utilize both visual and textual content when processing the keyframe. For example, the text information for keyframe 9 might be "sometime three years ago"; the text information for keyframe 10 might be "I still remember"; and if keyframe 7 does not contain any text, its text information could be empty.

[0119] In some embodiments, the server can use the identifier of the target video as a retrieval condition to extract object names, object reference images, and object relationship records from the object dataset associated with the target video, and output the extracted results as object reference information for the target video. It reads the text records associated with each current keyframe in the current keyframe sequence or extracts text content from each current keyframe, and outputs the text results of each current keyframe as text information for that current keyframe. The text information is arranged according to the chronological order in the current keyframe sequence, and current keyframes without text content are filled with empty text records, outputting a set of text information that corresponds one-to-one with the current keyframe sequence.

[0120] Step S1022: Fill the object reference information, text information, first description text, first identifier and second description text into the preset prompt word template to obtain the target prompt word.

[0121] A preset prompt template refers to a text organization framework determined before the current round of processing. The preset prompt template carries object reference information, text information, first descriptive text, first identifier, and second descriptive text. The purpose of the preset prompt template is to organize multiple types of input information into a unified text structure, enabling subsequent expressions of task requirements and input content in a fixed format. For example, a preset prompt template includes object reference information fields, text information fields, first descriptive text fields, first identifier fields, and second descriptive text fields; it may also include output format constraint fields to limit the arrangement of returned results. A target prompt is the complete prompt text formed after writing object reference information, text information, first descriptive text, first identifier, and second descriptive text into the preset prompt template. The target prompt is the text input carrier for the current round of processing. The purpose of the target prompt is to express the new input information corresponding to the current time window and the historical accumulated information together as structured text. For example, when the first identifier is [5], the second descriptive text includes “urban environment preparation” and “women enter the woods and gradually carry out their actions”, and the text information includes “sometime three years ago” corresponding to keyframe 9, a target prompt word containing all of the above can be generated.

[0122] In some embodiments, the server can read a preset prompt word template and locate the object reference information field, text information field, first description text field, first identifier field, and second description text field in the preset prompt word template, and output a populateable template instance. Then, the object reference information is written to the object reference information field, the text information is written to the text information field according to the time order of the current keyframe sequence, the first description text is written to the first description text field, the first identifier is written to the first identifier field, and the second description text is written to the second description text field. Finally, the written template instance is output as the target prompt word. The server can also perform a format check on the target prompt word to ensure that the fields are complete, the order is consistent, and the identifier format is valid, and output the text result that passes the check as the target prompt word for subsequent use.

[0123] Step S1023: Input the current keyframe sequence and target cue words into the pre-trained video understanding model.

[0124] A pre-trained video understanding model is a video processing model that has already completed parameter learning. It receives image and text data and performs joint understanding based on pre-defined parameter mappings. The purpose of a pre-trained video understanding model is to enable the simultaneous processing of visual content in the current keyframe sequence and textual constraints in target cues within a unified model.

[0125] In some embodiments, each current keyframe in the current keyframe sequence can be converted into an image input format supported by a pre-trained video understanding model, and the image results arranged in chronological order can be output as an image input sequence. Target prompts can be converted into a text input format supported by the pre-trained video understanding model, and the converted text results can be output as a text input sequence. Based on the input interface of the pre-trained video understanding model, the image input sequence and text input sequence are encapsulated into a joint input request, and the joint input request is sent to the pre-trained video understanding model.

[0126] Step S1024: Using the video understanding model, determine the third descriptive text, the second identifier, and the fourth descriptive text.

[0127] In some embodiments, see Figure 8 , Figure 8 This shows that step S1024 can be achieved by performing the following steps S301 to S303: Step S301: Generate initial results for the current keyframe sequence using the video understanding model. The initial results include initial frame description text for each current keyframe in the current keyframe sequence, initial identifiers of initial cumulative segmentation points corresponding to the cumulative keyframe sequence, and initial segment description text for each first video segment obtained after segmenting the cumulative keyframe sequence according to the initial cumulative segmentation points.

[0128] The initial results are a structured set of results generated after the video understanding model performs an initial analysis of the current keyframe sequence and target cues. These initial results precede subsequent calibration actions. They include keyframe-level results, segmentation location-level results, and video segment-level results. The purpose of the initial results is to establish a complete analysis foundation, which the server will then refine in subsequent steps. The initial frame description text is the first text description generated by the video understanding model for each current keyframe in the current keyframe sequence. At the data level, the initial frame description text belongs to keyframe-level text information. Its purpose is to provide a content representation for each current keyframe, providing a verification object for subsequent calibration actions. The initial cumulative segmentation points are the set of segmentation locations given by the video understanding model after the initial analysis of the cumulative keyframe sequence. These initial cumulative segmentation points are located on the cumulative keyframe sequence, not just on the current keyframe sequence. Their purpose is to provide preliminary segmentation boundaries for the cumulative keyframe sequence, allowing for further verification and adjustment in subsequent steps. The initial identifier is a location marker formed after recording the locations of the initial cumulative segmentation points. Initial identifiers can be represented by a list of keyframe numbers. The purpose of initial identifiers is to convert the initial cumulative segmentation points into location data that is easy for the server to read, store, and subsequently calibrate. The first video segment is a continuous segment formed after segmenting the cumulative keyframe sequence according to the initial cumulative segmentation points. The first video segment is a segment-level unit established based on the initial results. The purpose of the first video segment is to provide a clear segment scope for the initial segment description text. The initial segment description text is the segment-level text description generated by the video understanding model for each first video segment. The purpose of the initial segment description text is to provide the original segment representation for subsequent segment-level calibration.

[0129] For example, the current time window is keyframes 7 to 14, the historical keyframe sequence is keyframes 1 to 8, the first identifier is [5], and the second descriptive text includes "urban environment setup" and "the woman enters the woods and gradually unfolds her actions". After the server sends keyframes 7 to 14 and the target cue words to the video understanding model, it obtains the initial results. The initial frame descriptive text in the initial results can be: keyframe 7 is "the woman continues to walk in the woods", keyframe 8 is "the woman stops and looks ahead", keyframe 9 is "the woman speaks, and the subtitle shows some time three years ago", keyframe 10 is "the woman continues to speak in the woods", keyframe 11 is "the scene begins to suggest a memory transition", keyframe 12 is "the transition scene before the memory begins", keyframe 13 is "the scene enters the classroom, and the man and woman look at each other", keyframe 14 is "the classroom dialogue begins to unfold". The initial identifier in the initial results can be [5, 12, 13]. After the server segments the cumulative keyframe sequence from keyframe 1 to keyframe 14 according to [5,12,13], it obtains four first video segments. The initial segment description text of the four first video segments can be "Urban environment setup", "Woman moves in the forest and introduces memories through dialogue", "Transition before the memory begins", and "The scene enters the classroom and the memory plot begins".

[0130] Step S302: Based on the preset calibration text in the target prompt, calibrate the initial frame description text to obtain the third description text; calibrate the initial identifier to obtain the second identifier.

[0131] Preset calibration text is the pre-defined validation rule text within the target prompt. It constrains the expression and judgment methods of the initial results. Preset calibration text can include one or more of the following: keyframe description requirements, segmentation position review requirements, historical information inheritance requirements, text interference elimination requirements, and output format requirements. The purpose of preset calibration text is to provide the server with a unified validation basis, enabling the server to revise the initial frame description text and initial identifiers according to the same standard. For example, preset calibration text might require keyframe description text to include people, environment, and events, and require ignoring irrelevant text in the four corners of the screen; it might also require segmentation positions to remain consistent with historical results and require a second review of keyframes near the boundaries.

[0132] In some embodiments, calibrating the initial frame description text to obtain the third description text can be achieved as follows: First, for each current keyframe in the current keyframe sequence, obtain the current text corresponding to the current keyframe from the target prompt words; then, when the current text includes the first object identifier of the target object in the current keyframe, determine the second object identifier of the target object in the current keyframe from the initial frame description text; next, when the first object identifier is different from the second object identifier, replace the second object identifier in the initial frame description text with the first object identifier to obtain the third description text; when the current text does not include any first object identifier, or when the first object identifier is the same as the second object identifier, determine the initial frame description text as the third description text.

[0133] The current text is the text content associated with a current keyframe within the target prompt. The current text is set for a single current keyframe. It can reflect subtitles, names, or other text appearing near that keyframe. The purpose of the current text is to provide direct textual evidence for the initial object identification of the target object, enabling the server to have verifiable object name information at the keyframe level.

[0134] In some embodiments, after reading the target prompt and the current keyframe sequence, the server first locates the keyframe text segments in the target prompt according to the keyframe order in the current keyframe sequence, and outputs a candidate text set for each current keyframe. Then, it performs sequential matching and number matching on the candidate text set for each current keyframe, and outputs the matched text content as the current text corresponding to each current keyframe. Finally, it writes empty text records to the current keyframes for which no text content is matched, and outputs all the current texts in the order of the current keyframe sequence as the current text set.

[0135] The target object is the object in the current keyframe that requires object identification verification. The target object is not all objects in the current keyframe, but rather the objects that the current text has explicitly identified or can explicitly identify. The first object identifier is the identifier in the current text used to indicate the target object. The first object identifier can be a person's name or other markers that clearly point to the target object. The purpose of the first object identifier is to provide a baseline for object identification from the current text. The second object identifier is the identifier in the initial frame description text used to indicate the target object in the current keyframe. The second object identifier comes from the initial frame description text, not from the current text. The purpose of the second object identifier is to provide the object naming result in the initial frame description text for consistency verification with the first object identifier.

[0136] In some embodiments, the server can perform object identifier retrieval on the current text of each current keyframe and output a set of current keyframes including a first object identifier and a set of current keyframes excluding the first object identifier. First, for the set of current keyframes including the first object identifier, the object identifier content is extracted from the current text of each current keyframe, and the extracted object identifier content is output as the first object identifier of the target object in each current keyframe. Then, the initial frame description text corresponding to the set of current keyframes including the first object identifier is read, and after locating the object identifier content indicating the target object in each initial frame description text, a second set of object identifiers is output. Replacement refers to deleting the second object identifier used to indicate the target object in the initial frame description text and writing the first object identifier in the position of the second object identifier. Replacement is a targeted modification of the object identifier content and does not change other non-object identifier content in the initial frame description text. The purpose of replacement is to ensure that the target object name in the third description text is consistent with the object name in the current text.

[0137] In some embodiments, after reading the first object identifier, the second object identifier, and the initial frame description text, the text content of the first object identifier and the second object identifier are first compared, and the comparison result is output as an object identifier consistency result. Then, when the object identifier consistency result shows that the first object identifier and the second object identifier are different, the position of the second object identifier is located in the initial frame description text, and the second object identifier is replaced with the first object identifier, and the replaced text result is output. Finally, the replaced text result is written into the result set according to the keyframe order in the current keyframe sequence, and the result set is output as the corresponding entry in the third description text set. For example, if the initial frame description text is "Classroom dialogue begins, male looks at female A", the second object identifier is "female A", and the first object identifier is "female B", then "female A" is replaced with "female B" to obtain the third description text "Classroom dialogue begins, male looks at female B".

[0138] In some embodiments, after reading the current text, the first object identifier, the second object identifier, and the initial frame description text, it is first determined whether the current text includes any of the first object identifiers, and the determination result is output as the first conditional result. If the first conditional result indicates that the current text does not include any of the first object identifiers, the initial frame description text is directly written into the result set and output as the corresponding entry in the third description text set. If the first conditional result indicates that the current text includes the first object identifier, it continues to compare whether the first object identifier and the second object identifier are the same, and the comparison result is output as the second conditional result. If the second conditional result indicates that the first object identifier and the second object identifier are the same, the initial frame description text is directly written into the result set and output as the corresponding entry in the third description text set. For example, the current text of keyframe 9 is "sometime three years ago". The current text does not include any of the first object identifiers. Therefore, no replacement is performed, and the initial frame description text of keyframe 9, "A woman speaks, and the subtitle shows 'sometime three years ago'", is directly determined as the third description text of keyframe 9.

[0139] Here, after the initial frame description text is formed, the first object identifier in the current text is further used as the keyframe-level verification basis, and the first object identifier is directly compared with the second object identifier in the initial frame description text. When an inconsistency occurs, only the second object identifier is replaced. When the current text does not contain any first object identifier or the first object identifier is the same as the second object identifier, the initial frame description text remains unchanged. This gives the third description text a clear correction basis at the object naming level, which can improve the consistency of keyframe-level object identifiers and reduce text offset caused by unfounded rewriting, thereby enhancing the accuracy and stability of the third description text.

[0140] In some embodiments, calibrating the initial identifier to obtain a second identifier can be achieved in the following way: First, for any initial cumulative segmentation point, based on the third descriptive text and the first descriptive text in the target prompt, determine the semantic coherence between the preceding and following video segments corresponding to the initial cumulative segmentation point; then, when the semantic coherence indicates semantic coherence between the preceding and following video segments, delete the initial identifier of the initial cumulative segmentation point; when the semantic coherence indicates semantic incoherence between the preceding and following video segments, determine the initial identifier of the initial cumulative segmentation point as the second identifier.

[0141] Adjacent video segments are two consecutive video segments separated by an initial cumulative segmentation point. The video segment before the initial cumulative segmentation point is the preceding adjacent video segment, and the video segment after the initial cumulative segmentation point is the following adjacent video segment. Adjacent video segments share the same boundary position. The purpose of adjacent video segments is to provide a local semantic comparison range, enabling the server to analyze the content relationship between the two sides around a single initial cumulative segmentation point. Semantic coherence is the result data used to characterize the content continuity relationship between adjacent video segments. Semantic coherence reflects the degree of consistency between adjacent video segments in terms of characters, environment, events, and narrative progression. Semantic coherence can be represented using discrete results or hierarchical results. The purpose of semantic coherence is to provide a basis for determining whether to retain the initial cumulative segmentation point. For example, if both the preceding and following adjacent video segments depict a woman continuing her actions in the forest, and the narrative progresses along the same event chain, then semantic coherence can characterize semantic coherence; if the preceding adjacent video segment depicts a realistic scene in the forest, and the following adjacent video segment depicts a flashback scene in the classroom, then semantic coherence can characterize semantic incoherence.

[0142] In some embodiments, after reading the initial identifier, the third descriptive text, and the first descriptive text from the target prompt, an initial cumulative segmentation point is selected sequentially according to its position in the initial identifier. The currently selected initial cumulative segmentation point, the third descriptive text, and the first descriptive text are then output as the current determination data. Next, based on the current determination data, the segment ranges on both sides of the current initial cumulative segmentation point are located on the cumulative keyframe sequence. The third descriptive text and the first descriptive text within each segment range are then organized and output as semantic comparison data for adjacent video segments. Subsequently, character continuity analysis, environmental continuity analysis, and event continuity analysis are performed on the semantic comparison data for adjacent video segments. The analysis results are then summarized and output as the semantic coherence of the current initial cumulative segmentation point.

[0143] In some embodiments, when semantic coherence is represented by semantic coherence, the position record of the current initial cumulative segmentation point in the initial identifier is first located, and the location result is output as the position to be deleted. Then, the position record of the current initial cumulative segmentation point is removed from the initial identifier according to the position to be deleted, and the list of removed positions is output as the intermediate identifier result. The intermediate identifier result is written into the identifier update set of the current round, and the identifier retention result after deleting the current initial cumulative segmentation point is output.

[0144] In some embodiments, when semantic coherence represents semantic incoherence, the position records of the current initial cumulative segmentation points are first retained, and the retention results are output as candidate retained positions. Then, the candidate retained positions are added to the second identifier construction set, and after being arranged according to the order of the initial cumulative segmentation points in the cumulative keyframe sequence, they are output as a valid position in the second identifier. Finally, after all initial cumulative segmentation points have been processed, all valid positions in the second identifier construction set are summarized and output as the second identifier.

[0145] For example, the semantic coherence representation obtained for keyframe 5 is semantically incoherent because keyframes 1 to 4 depict the urban environment, while keyframes 5 to 12 depict the woman moving in the woods and evoking a memory through dialogue. There is a clear content break between adjacent video segments, so the identifier for keyframe 5 is retained. The semantic coherence representation obtained for keyframe 13 is also semantically incoherent because keyframe 12 shows a clearer transition to the memory, while keyframes 13 to 14 show the scene entering the classroom and the dialogue beginning. There are changes in environment and narrative level between adjacent video segments, so the identifier for keyframe 13 is retained. After processing all initial cumulative segmentation points, the retained keyframes 5 and 13 are combined to obtain the second identifier [5,13].

[0146] Here, a local semantic verification mechanism is established around any initial cumulative segmentation point. Instead of directly using the initial identifier, it combines the third descriptive text and the first descriptive text in the target prompt to determine the semantic coherence of the adjacent video segments on both sides of the initial cumulative segmentation point. When the semantics are coherent, the initial identifier is deleted, and when the semantics are not coherent, it is retained as the second identifier. This makes the retention and deletion of the segmentation position based on the semantic relationship between segments, which can reduce over-segmentation or incorrect retention caused by single-point judgment and enhance the consistency between the second identifier and the actual narrative boundary.

[0147] Step S303: The initial fragment description text is calibrated based on the second identifier to obtain the fourth description text.

[0148] In some embodiments, calibrating the initial segment description text based on the second identifier to obtain the fourth description text can be achieved in the following way: First, the accumulated keyframe sequence is segmented according to the initial accumulated segmentation point corresponding to the second identifier to obtain multiple second video segments; then, for each second video segment, when the second video segment includes multiple first video segments, the initial segment description texts of the multiple first video segments are semantically merged to obtain the fourth description text of the second video segment; when the second video segment includes one first video segment, the initial segment description text of the first video segment is determined as the fourth description text of the second video segment.

[0149] The second video segment is a continuous keyframe segment obtained by re-dividing the accumulated keyframe sequence based on the initial accumulated segmentation points preserved by the second identifier. The second video segment is based on the premise that the second identifier has been calibrated. Compared with the first video segment, the second video segment reflects the calibrated segmentation result. The purpose of the second video segment is to provide a segment range consistent with the second identifier for the subsequent fourth descriptive text.

[0150] In some embodiments, after reading the second identifier and the accumulated keyframe sequence, each position record in the second identifier is first mapped to the keyframe order of the accumulated keyframe sequence, and a set of segmentation positions is output; then, based on the set of segmentation positions, multiple consecutive intervals are divided in the accumulated keyframe sequence according to time sequence, and the keyframe set of each consecutive interval is output as a second video segment; finally, all second video segments are summarized in time sequence and output as a set of second video segments for subsequent processing of each second video segment.

[0151] Semantic merging refers to combining the initial descriptive texts of multiple first video segments within the same second video segment into a unified text result based on semantic continuity. Semantic merging is not a mechanical splicing or deletion of any text; rather, it's about forming a single expression for the second video segment based on the content relationships of multiple initial descriptive texts. The purpose of semantic merging is to ensure that the fourth descriptive text covers the entire content range of the second video segment. For example, the initial descriptive texts for keyframes 5 to 11 are "A woman moves in the woods and evokes a memory through dialogue," while the initial descriptive text for keyframe 12 is "A transition before the memory begins." After semantically merging these two texts, the fourth descriptive text "A woman moves in the woods and evokes a memory through dialogue" can be formed.

[0152] In some embodiments, after reading the output set of second video segments and the initial segment description text of all first video segments, the number of first video segments covered by each second video segment is first determined, and the second video segments covered by more than one first video segment are output as segments to be semantically merged. For each segment to be semantically merged, the initial segment description text of multiple first video segments within the segment range is extracted, and semantic merging is performed on the multiple initial segment description texts to output a single merged text. Finally, the single merged text is written into the result set according to the time order of the second video segments and output as the fourth description text of the second video segment. The second video segments covered by one first video segment are output as directly determined segments. The initial segment description text of the first video segment is extracted from a first video segment associated with each directly determined segment, and the extracted text is directly written into the result set. Finally, the text in the result set is summarized according to the time order of the second video segments and output as the fourth description text of that part of the second video segment.

[0153] Here, after the second identifier has been formed, the second video segment is redefined using the second identifier, and two cases are distinguished: the second video segment includes multiple first video segments and the second video segment includes one first video segment. In the former case, the initial segment description text of multiple first video segments is semantically merged, and in the latter case, the initial segment description text of one first video segment is directly inherited. This ensures that the fourth description text is consistent with the segment boundary defined by the second identifier, reduces text mismatch, duplicate descriptions, and segment content omissions caused by changes in the segment range after calibration, and improves the boundary consistency and segment-level expression completeness of the fourth description text.

[0154] Here, a calibration path is introduced after the video understanding model outputs. The server first generates initial results covering keyframe level, segmentation position level, and video segment level. Then, it uses preset calibration text to perform targeted revisions on the initial frame description text and initial identifier. Finally, it adjusts the initial segment description text for consistency based on the second identifier. This creates a hierarchical constraint relationship between keyframe-level text, segmentation position, and segment-level text, reduces the propagation of local deviations caused by one-time generation, enhances the consistency between segmentation boundaries and text expression, and improves the matching degree between the fourth description text and the second identifier.

[0155] Here, by introducing object reference information and text information in the current keyframe sequence processing stage, and uniformly filling the object reference information, text information, first descriptive text, first identifier, and second descriptive text into a preset prompt word template to form target prompt words, the current keyframe sequence and target prompt words are then input into a pre-trained video understanding model. This allows the server to jointly utilize visual content, text content, and historical results to generate third descriptive text, second identifier, and fourth descriptive text in the same processing chain. Therefore, it can enhance the semantic expression integrity of the current keyframe sequence, improve the consistency of object recognition, enhance the utilization of contextual information when determining the segmentation position, and simultaneously achieve the integrated output of keyframe-level results, segmentation identifier results, and video segment-level results, thereby improving the stability, accuracy, and processing efficiency of video segmentation-related results.

[0156] Step S103: Based on the third description text, the second identifier, and the fourth description text, update the cumulative state information to obtain the cumulative state information of the next time window of the current time window.

[0157] The next time window is the time window that immediately follows the current time window to enter the next iteration of processing. The current time window and the next time window partially overlap in the time dimension. The time dimension represents the position of video data in terms of chronological order, and the position in the time dimension can be represented by the keyframe time position. Partial overlap means that the current time window and the next time window share a common time range, but they are not exactly the same. The purpose of partial overlap is to allow keyframes near the boundary to be observed again when entering the next time window.

[0158] In some embodiments, the server writes a third descriptive text into the keyframe description portion of the accumulated state information to obtain an updated keyframe description portion. It then writes a second identifier into the segmentation identifier portion of the accumulated state information to obtain an updated segmentation identifier portion. Finally, it writes a fourth descriptive text into the video segment description portion of the accumulated state information to obtain an updated video segment description portion. Finally, the updated keyframe description portion, the updated segmentation identifier portion, and the updated video segment description portion are merged to obtain the accumulated state information for the next time window.

[0159] In some embodiments, the cumulative state information is updated based on the third description text, the second identifier, and the fourth description text to obtain the cumulative state information of the next time window of the current time window. This can be achieved in the following way: First, the third description text is concatenated after the first description text in the cumulative state information to obtain the merged description text; then, the merged description text is determined as the first description text in the cumulative state information of the next time window, the second identifier is determined as the first identifier in the cumulative state information of the next time window, and the fourth description text is determined as the second description text in the cumulative state information of the next time window.

[0160] The merged description text is the text result obtained by concatenating the first description text from the accumulated state information with the third description text formed in the current time window, according to a predetermined order. The merged description text still belongs to the keyframe-level text set, but it simultaneously inherits text records from both the historical keyframe sequence and the current keyframe sequence. The purpose of the merged description text is to incorporate the keyframe-level text results produced in the current time window into existing keyframe-level text results, allowing the next time window to directly read continuous keyframe-level text information without needing to re-retrace historical processing results. For example, when the current time window is keyframes 7 to 14, the first description text in the accumulated state information includes text records from keyframes 1 to 8, and the third description text includes text records from keyframes 7 to 14. Concatenating these two parts of the description text yields the merged description text. Furthermore, the first description text for keyframes 1 to 8 is placed first, followed by the third description text for keyframes 7 to 14.

[0161] In some embodiments, after reading the first description text in the accumulated state information and the third description text obtained in the current time window, the two description texts are joined together in the order of the first description text first and the third description text last to form a continuous text sequence. Finally, the continuous text sequence is sorted and the record integrity is checked before the merged description text is output. The merged description text serves as the basis for updating the first description text in the accumulated state information of the next time window.

[0162] In some embodiments, the merged description text is written into the first description text field of the cumulative status information of the next time window, and the updated first description text is output; the second identifier is written into the first identifier field of the cumulative status information of the next time window, and the updated first identifier is output; the fourth description text is written into the second description text field of the cumulative status information of the next time window, and the updated second description text is output; finally, the updated first description text, the updated first identifier, and the updated second description text are summarized and output as the cumulative status information of the next time window.

[0163] Here, the cumulative state information is updated using a clear data write-back path. The server first concatenates the third description text to the existing first description text to form a merged description text. Then, the merged description text, the second identifier, and the fourth description text are written into the first description text, the first identifier, and the second description text in the cumulative state information of the next time window, respectively. This ensures that the keyframe-level text results, segmentation position results, and video segment-level text results are continuously transmitted between time windows, reducing information loss and field misalignment caused by state reconstruction during iteration, and improving data integrity and state consistency when reading historical results in the next time window.

[0164] In some embodiments, the processing range continues to move along the time dimension according to a preset time window, and retains the portion of the interval that overlaps with the current time window to obtain the next time window. The accumulated state information of the next time window and the next time window are incorporated into the subsequent processing flow to obtain the preparation state for the next round of iteration processing.

[0165] In some embodiments, after updating the accumulated state information, it can be determined whether the current time window is the last time window of the preset time window, that is, whether the iterative processing of all keyframes in the target video has been completed. If the iterative processing of all keyframes in the target video has been completed, then step S104 is executed; if the iterative processing of all keyframes in the target video has not been completed, then the next time window is taken as the new current time window, and the process jumps to step S101 to continue to obtain the current keyframe sequence and accumulated state information of the target video in the new current time window, and a new round of iterative processing is started.

[0166] Step S104: After completing the iterative processing of all keyframes in the target video, the target video is segmented based on the accumulated state information obtained after the iterative processing to obtain the video segmentation result of the target video.

[0167] All keyframes refer to the set of all keyframes in the target video within the processing scope of this application embodiment. All keyframes completing iterative processing means that there are no keyframes in the target video that have not yet entered the iterative processing flow. For example, if the target video has 60 keyframes, keyframes 1 to 60 have all entered the iterative processing flow. Video segmentation is the process of dividing the target video into multiple consecutive video segments based on the segmentation positions in the accumulated state information. The video segmentation result is the segmentation result data output based on the accumulated state information obtained after iterative processing. The video segmentation result at least reflects the boundaries of the multiple video segments after the target video has been divided, and may also carry descriptive text for each video segment.

[0168] In some embodiments, the server checks whether the results of each iteration have covered all keyframes in the target video, and determines the end of the iteration process when all keyframes have been processed. After the iteration process ends, the final retained cumulative state information is read, i.e., the cumulative state information obtained after the iteration process. Based on the segmentation positions in the cumulative state information obtained after the iteration process, the target video is segmented to obtain multiple video segments. The multiple video segments and the text records in the cumulative state information are combined and organized to obtain the video segmentation result of the target video.

[0169] In some embodiments, see Figure 9 , Figure 9 Step S104 can be achieved by performing the following steps S1041 to S1043: Step S1041: Obtain the identifier sequence corresponding to the identifier of all segmentation points of the target video from the accumulated state information obtained after iterative processing.

[0170] All segmentation points refer to the set of all segmentation locations confirmed in the target video after all iterations of processing. All segmentation points define all boundary locations when the target video is divided into multiple parts. All segmentation points can be represented by keyframe numbers or by temporal positions. For example, in the same example, keyframes 5, 13, and 17 are confirmed as all segmentation points after all iterations. The identifier sequence refers to the set of identifiers for all segmentation points arranged in chronological order. Each identifier in the identifier sequence represents a segmentation location, and the order in the identifier sequence indicates the execution order of video segmentation. The identifier sequence is ordered, allowing subsequent processing to be performed directly based on it. For example, when all segmentation points are keyframes 5, 13, and 17, the identifier sequence could be [5, 13, 17].

[0171] In some embodiments, the cumulative state information obtained after iterative processing is used as input. The identifiers of all segmentation points for the target video are read from the cumulative state information, and the identifiers of all segmentation points are arranged in chronological order to output an identifier sequence. For example, in a target video with a duration of 120 seconds, after completing all iterative processing of keyframes 1 to 20, the identifiers 5, 13, and 17 of all segmentation points are read from the cumulative state information and arranged in chronological order of the keyframe numbers to form an identifier sequence [5, 13, 17].

[0172] Step S1042: Based on the identifiers in the identifier sequence, the target video is segmented sequentially to obtain multiple target video segments.

[0173] A target video segment refers to a continuous video portion formed after segmenting a target video based on identifiers in an identifier sequence. Each target video segment has a clearly defined start and end position, and adjacent target video segments are separated by a segmentation position. Multiple target video segments together constitute the segmented structure of the target video. For example, keyframes 1 to 4 can form one target video segment, and keyframes 5 to 12 can form another target video segment.

[0174] In some embodiments, the server takes an identifier sequence as input, determines the segmentation position in the target video one by one according to the order of the identifiers in the identifier sequence, performs video segmentation processing one by one, and outputs multiple target video segments.

[0175] In some embodiments, the target video is segmented sequentially based on the identifiers in the identifier sequence to obtain multiple target video segments. This can be achieved as follows: First, the timestamp of the segmentation point corresponding to each identifier in the identifier sequence is obtained from a preset mapping table; then, the start and end time range of each target video segment in the target video is determined based on the multiple timestamps arranged in chronological order; finally, the target video is segmented according to each start and end time range to obtain multiple target video segments.

[0176] Step S1043: Determine the video segmentation result of the target video from multiple target video segments.

[0177] Using multiple target video segments as input, the results of the multiple target video segments are sorted in chronological order, and the sorted multiple target video segments are determined as the video segmentation results of the target video.

[0178] Here, based on the accumulated state information obtained after iterative processing, the identifiers of all segmentation points for the target video are first obtained to form an identifier sequence. Then, the target video is segmented sequentially based on the identifiers in the identifier sequence. Finally, multiple target video segments are directly determined as the video segmentation results of the target video. This makes the segmentation information in the accumulated state information and the actual video segmentation action a clear and consistent relationship, which can reduce the inconsistency between the segmentation position expression and the segmentation execution, ensure clear segmentation order, clear segmentation boundaries and stable result structure, reduce the additional processing burden in the result generation stage, and improve the accuracy, executability and processing efficiency of the video segmentation results of the target video.

[0179] In some embodiments, after obtaining the video segmentation result of the target video, the following steps may also be performed: First, obtain the fifth description text of each target video segment from the accumulated state information obtained after iterative processing; then, generate a segment summary of each target video segment based on the fifth description text; finally, generate plot index information of the target video based on the start and end time range and segment summary corresponding to each target video segment; the plot index information is used to provide a segment summary display of the corresponding target video segment when playing the target video.

[0180] In this embodiment, by performing iterative processing on the keyframes of the target video according to a preset time window, and by partially overlapping the current time window with the next time window in the time dimension, the computational load of processing all keyframes in the target video at once is effectively reduced, while ensuring the continuity of information in the time dimension. During iterative processing, accumulated state information is obtained, including a first descriptive text for each historical keyframe in the historical keyframe sequence preceding the current time window, a first identifier of the historical segmentation point corresponding to the historical keyframe sequence, and a second descriptive text for each video segment obtained after segmenting the historical keyframe sequence according to the historical segmentation point. This accumulated state information is combined with the current keyframe sequence of the target video within the current time window. Based on the current keyframe sequence, the first descriptive text, the first identifier, and the second descriptive text, a semantic association between the current keyframe sequence and the historical keyframe sequence can be established. This accurately determines the third descriptive text for each current keyframe in the current keyframe sequence, the second identifier of the cumulative segmentation point corresponding to the cumulative keyframe sequence, and the fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation point. Furthermore, the accumulated state information is updated based on the third descriptive text, the second identifier, and the fourth descriptive text to obtain the accumulated state information for the next time window of the current time window. This allows the iterative processing of keyframes to dynamically adjust the accumulated state information based on the continuously acquired current keyframe sequence, thereby automatically correcting misjudgments or descriptive deviations in segmentation points caused by insufficient local keyframe information in the early stages. Finally, when completing the iterative processing of all keyframes in the target video, the target video is segmented based on the continuously updated accumulated state information obtained after the iterative processing. This effectively avoids incorrect video segmentation caused by sudden changes in local scenes or complex editing techniques, significantly improving the accuracy of the video segmentation results of the target video, as well as enhancing the logical coherence and semantic integrity of the descriptive text corresponding to each video segment.

[0181] In some embodiments, prior to any of the above embodiments, a video understanding model training method may also be provided. The video understanding model trained using this method can be applied to any of the above video segmentation methods. This video understanding model training method can be executed by a video understanding model training device. This training device can be the same electronic device as the electronic device used to implement the video segmentation method, or it can be a different electronic device. That is, the video understanding model training device used to implement the video understanding model training method and the video segmentation device used to implement the video segmentation method can be located in the same electronic device or in different electronic devices. The video understanding model training method of this application embodiment can be executed by a server, by a terminal, or through interaction between a server and a terminal.

[0182] See Figure 10 , Figure 10 This is a flowchart illustrating the training method of the video understanding model provided in this application embodiment, which will be combined with... Figure 10 The steps shown are explained below, taking the server as the execution subject of the video understanding model training method as an example. The training method of this video understanding model, through steps S401 to S406, performs the following iterative processing on the sample keyframes of the sample videos in the training sample set according to a preset training time window: Step S401: Obtain the current keyframe sequence, cumulative state information, and target output truth value of the sample video within the current training time window. The cumulative state information includes: the first description text of each historical keyframe in the historical keyframe sequence before the current training time window, the first identifier of the historical segmentation point corresponding to the historical keyframe sequence, and the second description text of each sample video segment obtained after segmentation according to the first identifier. The target output truth value includes: the frame description truth value sequence of the sample keyframes of the sample video, the identifier truth value sequence of the actual segmentation point of the sample video, and the segment description truth value sequence of the sample video segments obtained after segmenting the sample video according to the actual segmentation point.

[0183] The sample video is any training video in the training sample set. The sample video can be a manually annotated segment from a TV series, movie, or online video. The current training time window is the time range covered by the current iteration during training. The current keyframe sequence of the sample is the set of sample keyframes located within the current training time window and arranged chronologically. The cumulative state information of the sample is the historical training context data retained before entering the current training time window. The first description text of the sample is the manually annotated description text for each historical keyframe in the sample's historical keyframe sequence. The first identifier of the sample is the manually annotated location marker for the historical segmentation point of the sample. The second description text of the sample is the manually annotated description text for each sample video segment obtained after segmenting the historical keyframe sequence according to the first identifier of the sample. The target output ground truth is a pre-annotated set of global standard answers used to supervise the training of the video understanding model. Since iterative processing across time windows is involved for the sample video, the target output ground truth covers multiple ground truth sequences for the entire sample video. The target output ground truth includes at least: a frame description ground truth sequence, an identifier ground truth sequence, and a segment description ground truth sequence. The frame description truth sequence refers to the pre-annotated truth text sequence of sample keyframes in the sample video. The frame description truth sequence is arranged chronologically according to the sample keyframes and is used to provide frame-level supervision information. It contains the standard answers to the content descriptions of the sample keyframes throughout the entire playback lifecycle of the sample video. The identifier truth sequence refers to the pre-annotated truth identifier sequence of the sample video's true segmentation points. The identifier truth sequence is arranged chronologically according to the sample keyframes or segmentation point positions and is used to indicate whether a true segmentation point exists at each position. It contains the standard answers to the locations of the true segmentation points in the sample video. The segment description truth sequence refers to the truth text sequence of the sample video segments formed after segmenting the sample video according to the true segmentation points. The segment description truth sequence is arranged chronologically according to the sample video segments and is used to provide segment-level supervision information. It contains the standard answers to the semantic descriptions of the sample video segments obtained after segmenting the sample video according to the aforementioned true segmentation points. In some embodiments, when the current training time window is the first training time window, the sample cumulative state information is initialized to empty.

[0184] In some embodiments, taking a 120-second sample video as an example, the sample keyframes are numbered in chronological order as sample keyframe 1 to sample keyframe 20, and the training time windows are sample keyframe 1 to sample keyframe 8, sample keyframe 7 to sample keyframe 14, and sample keyframe 13 to sample keyframe 20 in sequence. When the current training time window is sample keyframe 7 to sample keyframe 14, the server obtains the current keyframe sequence of the sample keyframes sample keyframes 7 to 14, obtains the first description text and the first identifier of the sample corresponding to the historical keyframe sequence of the sample keyframes sample keyframes 1 to 8 [5], and obtains the second description text of the sample video segments formed after being segmented according to the first identifier of the sample [5]. At the same time, the target output truth value corresponding to the current training time window is obtained. The identifier truth value can be [5, 13], the frame description truth value can be the manually annotated description text of each of the sample keyframes 7 to 14, and the segment description truth value can be the manually annotated description text of each sample video segment formed after being segmented according to the identifier truth value.

[0185] Step S402: Noise injection processing is performed on the first description text of the sample to obtain noisy first description text; noise injection processing is performed on the first identifier of the sample to obtain noisy first identifier; noise injection processing is performed on the second description text of the sample to obtain noisy second description text.

[0186] Noise injection is used to proactively construct input states with historical errors during the training phase, enabling the video understanding model to learn and recover from inaccurate historical contexts. For the sample first description text, the server can randomly select a portion of the sample first description text and perform object identifier replacement, event expression perturbation, or local word deletion on the selected sample first description text to obtain noisy first description text. For the sample first identifier, the server can randomly delete a portion of the sample first identifier, add sample first identifiers at incorrect positions, or move the original sample first identifier to the historical keyframe position of an adjacent sample to obtain noisy first identifier. For the sample second description text, the server can randomly replace the object identifiers or local event descriptions in the sample second description text to obtain noisy second description text.

[0187] In some embodiments, the above noise injection processing is performed using a random strategy, which can be controlled by a preset probability to determine the injection ratio of various types of noise. For example, in the first sample description text "woman enters the forest", the server replaces "woman" with "man" to obtain the noisy first description text "man enters the forest"; in the first sample identifier [5], the server modifies position 5 to position 4 to obtain the noisy first identifier [4]; in the second sample description text "woman enters the forest and gradually unfolds her actions", the server replaces "woman" with "man" to obtain the noisy second description text "man enters the forest and gradually unfolds his actions". Through this processing, the historical context received by the video understanding model to be trained during training is no longer always in an ideal state, which is beneficial to improving the error correction and fault tolerance capabilities of the video understanding model to be trained in actual application scenarios.

[0188] Step S403: Based on the current keyframe sequence of the sample, the noisy first descriptive text, the noisy first identifier, and the noisy second descriptive text, the prediction result is determined through the video understanding model to be trained. The prediction result includes: the predicted third descriptive text for each current keyframe in the current keyframe sequence of the sample, the predicted second identifier for the cumulative segmentation point of the sample cumulative keyframe sequence, and the predicted fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence of the sample according to the cumulative segmentation point. The cumulative keyframe sequence of the sample includes the current keyframe sequence of the sample and the historical keyframe sequence of the sample.

[0189] The video understanding model to be trained is either not yet fully trained or still undergoing iterative optimization. The server uses the current keyframe sequence of the sample, along with noisy first descriptive text, noisy first identifier, and noisy second descriptive text, as input to the model, enabling the video understanding model to output prediction results within a noisy historical context. The predicted third descriptive text represents the keyframe-level description result given by the video understanding model to each current keyframe in the current keyframe sequence of the sample. The predicted second identifier represents the segmentation position prediction result given by the video understanding model to the cumulative keyframe sequence of the sample. The predicted fourth descriptive text represents the segment-level description result given by the video understanding model to the sample video segments formed after segmentation according to the cumulative segmentation points of the sample.

[0190] In some embodiments, taking the above-mentioned sample keyframes 1 to 20 as examples, when the current training time window is sample keyframes 7 to 14, the server inputs sample keyframes 7 to 14, noisy first descriptive text, noisy first identifier [4], and noisy second descriptive text into the video understanding model to be trained. The video understanding model to be trained outputs predicted third descriptive text, for example, for sample keyframe 7, it outputs "The woman continues to walk in the forest", and for sample keyframe 13, it outputs "The scene enters the classroom, and the man and woman make eye contact"; at the same time, it outputs predicted second identifier, for example [5,13]; and outputs predicted fourth descriptive text, for example, "Urban environment setup", "The woman moves in the forest and evokes memories through dialogue", "The scene enters the classroom, and the memory plot begins". The above prediction results are then entered into the loss calculation process together with the target output ground truth.

[0191] Step S404: Based on the predicted third description text and frame description truth sequence, the predicted second identifier and identifier truth sequence, and the predicted fourth description text and fragment description truth sequence, the loss is calculated to obtain the multi-task joint loss value.

[0192] The multi-task joint loss value is used to comprehensively reflect the prediction errors of the video understanding model under training on keyframe description, video segmentation, fragment description, and overall output tasks. The server calculates the keyframe description loss value based on the difference between the predicted third description text and the ground truth sequence of frame descriptions; the video segmentation loss value based on the difference between the predicted second label and the label ground truth sequence; the fragment description loss value based on the difference between the predicted fourth description text and the fragment description ground truth sequence; and further calculates the global loss value based on the overall difference between the prediction results and the target output ground truth. Finally, the multiple loss terms are jointly weighted to obtain the multi-task joint loss value. In this way, the video understanding model under training does not only focus on a single segmentation task during training, but also simultaneously optimizes keyframe-level semantic understanding, segmentation point localization, fragment-level semantic generalization, and overall output consistency.

[0193] In some embodiments, the loss calculation based on the predicted third descriptive text and frame description truth sequence, the predicted second identifier and identifier truth sequence, and the predicted fourth descriptive text and fragment description truth sequence to obtain a multi-task joint loss value can be achieved in the following way: First, from the frame description truth sequence, obtain the target frame description truth value for the current keyframe of each sample; from the identifier truth sequence, obtain the target identifier truth value of the target true segmentation point corresponding to the accumulated keyframe sequence of the samples; from the fragment description truth sequence, obtain the target fragment description truth value of each sample video fragment obtained after segmenting the accumulated keyframe sequence of the samples according to the target true segmentation point; then, based on the predicted third descriptive text and target frame description truth value, the predicted second identifier and target identifier truth value, and the predicted fourth descriptive text and target fragment description truth value, the loss calculation is performed to obtain a multi-task joint loss value.

[0194] The target frame description ground value refers to the set of true descriptive texts selected from the frame description ground value sequence that match the current keyframe of each sample one by one according to its frame position. The target frame description ground value is used to constrain the prediction of the third descriptive text. The server receives the frame description ground value sequence and the current keyframe of the sample, performs localization filtering in the frame description ground value sequence according to the frame position of each current keyframe in the sample video, arranges the filtered true descriptive texts in the order of the current keyframes of the sample, and outputs the target frame description ground value.

[0195] The target true segmentation point refers to the true segmentation point located within the coverage area of ​​the accumulated keyframe sequence. The target true segmentation point defines the true boundary position of the segmentation supervision within the current training time window. The target label ground truth is the set of true labels extracted from the label ground truth sequence, used to indicate the distribution of the target true segmentation points. The target label ground truth is used to constrain the prediction of the second label. The server receives the label ground truth sequence and the accumulated keyframe sequence. First, based on the start and end frame positions of the accumulated keyframe sequence, it determines the value range of the target true segmentation point. Then, it extracts the true labels falling within this value range from the label ground truth sequence, arranges them according to the chronological order of the accumulated keyframe sequence, and outputs the target label ground truth.

[0196] The target fragment description ground truth refers to the set of true descriptive texts extracted from the fragment description ground truth sequence for each sample video fragment corresponding to the sample accumulated keyframe sequence. The target fragment description ground truth is used to constrain the prediction of the fourth descriptive text. The server receives the fragment description ground truth sequence, the sample accumulated keyframe sequence, and the target true segmentation points. First, according to the position of the target true segmentation points in the sample accumulated keyframe sequence, the server divides the sample accumulated keyframe sequence into multiple sample video fragments. Then, according to the chronological order of the sample video fragments, the server selects the true descriptive text for each sample video fragment from the fragment description ground truth sequence and outputs the target fragment description ground truth.

[0197] In some embodiments, see Figure 11 , Figure 11 The method demonstrates how to calculate the joint loss value based on the predicted ground truth values ​​of the third descriptive text and the target frame description, the predicted ground truth values ​​of the second identifier and the target identifier, and the predicted ground truth values ​​of the fourth descriptive text and the target fragment description. This can be achieved by executing the following steps S4041 to S4045: Step S4041: Calculate the first loss based on the predicted third descriptive text and the target frame descriptive ground truth to obtain the keyframe descriptive loss value.

[0198] The keyframe description loss value is used to characterize the prediction error of the video understanding model to be trained on the keyframe-level description task.

[0199] In some embodiments, the first loss calculation based on the predicted third descriptive text and the target frame descriptive ground value to obtain the keyframe descriptive loss value can be achieved in the following way: First, the one-hot encodings of the positions of each word in the target frame descriptive ground value in the preset dictionary are merged into a first one-hot encoding sequence; the first prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting the predicted third descriptive text is obtained; then, the first cross-entropy loss value is determined based on the first prediction probability in the first prediction probability distribution matrix and the first one-hot encoding label in the first one-hot encoding sequence; finally, the first cross-entropy loss value is determined as the keyframe descriptive loss value.

[0200] The pre-defined dictionary is a set of words established beforehand during the training phase. It includes words used in training, punctuation marks, and pre-defined special markers. One-hot encoding refers to an encoding format where only the target word position in the pre-defined dictionary is set to 1, and all other word positions are set to 0. The first one-hot encoded sequence is an encoding sequence formed by concatenating the one-hot encodings of each word in the target frame description ground truth in output order. The first predicted probability distribution matrix is ​​the set of probability distribution results generated by the video understanding model for each output position of all words in the pre-defined dictionary when generating predicted third-party descriptive text. The first predicted probability is the predicted probability value of any word corresponding to any output position in the first predicted probability distribution matrix. The first one-hot encoded label is the label value used in the first one-hot encoded sequence to indicate the position of the target word. The first cross-entropy loss value is the loss result calculated based on the difference between the first predicted probability and the first one-hot encoded label. The keyframe description loss value is used to characterize the degree of difference between the predicted third-party descriptive text and the target frame description ground truth.

[0201] In some embodiments, the predicted third descriptive text, the target frame description ground value, and a preset dictionary are used as inputs. First, each word in the target frame description ground value is converted into a first one-hot encoded sequence. Then, the first prediction probability distribution matrix of the predicted third descriptive text at each output position is read from the video understanding model output. Next, cross-entropy loss is calculated based on the first prediction probability in the first prediction probability distribution matrix and the first one-hot encoded label in the first one-hot encoded sequence. The first cross-entropy loss value is output and determined as the keyframe description loss value.

[0202] Here, each row of the first predicted probability distribution matrix represents the probability distribution given by the video understanding model to be trained for all words in the preset dictionary at an output word position. The first one-hot encoded label represents the target word position of the target frame description ground value at that output word position. The first cross-entropy loss value is used to measure the degree of difference between the predicted third descriptive text and the target frame description ground value at the word generation level. By calculating the keyframe description loss value based on the predicted third descriptive text and the target frame description ground value, the keyframe-level semantic expression ability of the video understanding model can be directly constrained during the training process. This allows the video understanding model to gradually learn the accurate description methods of people, scenes, and events in the current keyframe sequence of the sample, thereby improving the consistency between the predicted third descriptive text and the target frame description ground value, and thus improving the accuracy and stability of the trained video understanding model in recognizing keyframe content.

[0203] Step S4042: Calculate the second loss based on the predicted second identifier and the true value of the target identifier to obtain the video segmentation loss value.

[0204] The video segmentation loss value is used to characterize the prediction error of the video understanding model to be trained on the task of locating segmentation points in the sample accumulation.

[0205] In some embodiments, the video segmentation loss value is obtained by calculating the second loss based on the predicted second identifier and the ground truth value of the target identifier. This can be achieved in the following way: First, for the predicted second identifier of each sample cumulative segmentation point corresponding to the sample cumulative keyframe sequence, the cross-entropy loss is calculated based on the predicted second identifier and the ground truth value of the target identifier to obtain the cross-entropy loss value; then, the cross-entropy loss values ​​of all sample cumulative segmentation points corresponding to the sample cumulative keyframe sequence are aggregated to obtain the video segmentation loss value.

[0206] The predicted second identifier is the segmentation position prediction result output by the video understanding model for the accumulated keyframe sequence. The accumulated segmentation point is the candidate segmentation position in the accumulated keyframe sequence. The cross-entropy loss value is the loss result calculated for any accumulated segmentation point position based on the difference between the predicted second identifier and the ground truth target identifier. Aggregation processing is a method of summing up the cross-entropy loss values ​​at multiple accumulated segmentation point positions. The video segmentation loss value is used to characterize the overall degree of difference between the predicted second identifier and the ground truth target identifier in the segmentation point recognition task.

[0207] In some embodiments, the predicted second identifier, the target identifier ground value, and all sample cumulative segmentation points in the sample cumulative keyframe sequence are used as input. For each sample cumulative segmentation point, cross-entropy loss is calculated based on the predicted second identifier and the target identifier ground value to obtain multiple cross-entropy loss values. Then, the multiple cross-entropy loss values ​​are aggregated to output the video segmentation loss value.

[0208] Here, cross-entropy loss is used to measure the difference between the predicted second identifier and the ground truth target identifier at any cumulative segmentation point position. Aggregation processing can be summation, averaging, or weighted summation. The smaller the video segmentation loss value, the closer the predicted second identifier is to the ground truth target identifier. By setting the video segmentation loss value, the video understanding model can be constrained to improve its segmentation point localization ability in the cumulative keyframe sequence. By calculating the video segmentation loss value based on the predicted second identifier and the ground truth target identifier, the cumulative segmentation point recognition ability of the video understanding model can be directly constrained during training. This allows the video understanding model to gradually learn the true segmentation boundary distribution in the cumulative keyframe sequence, thereby reducing the deviation between the predicted second identifier and the ground truth target identifier, and ultimately improving the accuracy of the trained video understanding model in locating segmentation points and the reliability of the video segmentation results.

[0209] Step S4043: Calculate the third loss based on the predicted fourth descriptive text and the target fragment description ground truth value to obtain the fragment description loss value.

[0210] The segment description loss value is used to characterize the prediction error of the video understanding model to be trained on the video segment-level description task.

[0211] In some embodiments, the fragment description loss value is obtained by calculating the third loss based on the predicted fourth descriptive text and the target fragment description ground value. This can be achieved as follows: First, the one-hot codes of the positions of each word in the target fragment description ground value in the preset dictionary are merged into a second one-hot coding sequence; the second prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting the fourth descriptive text is obtained; then, the second cross-entropy loss value is determined based on the second prediction probability in the second prediction probability distribution matrix and the second one-hot coding label in the second one-hot coding sequence; finally, the second cross-entropy loss value is determined as the fragment description loss value.

[0212] The second one-hot encoding sequence is the encoding sequence formed by concatenating the one-hot encodings of each word in the ground truth description of the target fragment in the output order. The second prediction probability distribution matrix is ​​the set of probability distribution results generated by the video understanding model for each output position of all words in the preset dictionary when generating the predicted fourth descriptive text. The second prediction probability is the predicted probability value of any word corresponding to any output position in the second prediction probability distribution matrix. The second one-hot encoding label is the label value in the second one-hot encoding sequence used to indicate the position of the target word. The second cross-entropy loss value is the loss result calculated based on the difference between the second prediction probability and the second one-hot encoding label. The fragment description loss value is used to characterize the degree of difference between the predicted fourth descriptive text and the ground truth description of the target fragment.

[0213] In some embodiments, the server takes the predicted fourth descriptive text, the target fragment description ground value, and a preset dictionary as input. First, it converts each word in the target fragment description ground value into a second one-hot encoded sequence. Then, it reads the second prediction probability distribution matrix of the predicted fourth descriptive text at each output position from the video understanding model output. Next, it performs cross-entropy loss calculation based on the second prediction probability in the second prediction probability distribution matrix and the second one-hot encoded label in the second one-hot encoded sequence, outputs the second cross-entropy loss value, and determines the second cross-entropy loss value as the fragment description loss value.

[0214] Each row in the second prediction probability distribution matrix represents the prediction probability distribution of the fourth descriptive text at an output position for all words in the preset dictionary. The second one-hot encoded label represents the target word position of the target fragment description truth at that output position. The smaller the second cross-entropy loss value, the closer the predicted fourth descriptive text is to the target fragment description truth. By setting the fragment description loss value, the video understanding model can be constrained to improve the semantic generalization ability of sample video fragments.

[0215] Here, by calculating the segment description loss value based on the predicted fourth description text and the target segment description ground value, the segment-level semantic generalization ability of the video understanding model can be directly constrained during the training process. This allows the video understanding model to gradually learn how to summarize the overall semantics of sample video segments according to the segmentation results, thereby improving the consistency between the predicted fourth description text and the target segment description ground value. In turn, this improves the accuracy, completeness, and semantic differentiation ability of the trained video understanding model in summarizing video segment content.

[0216] Step S4044: Based on the prediction results and the target output true value, perform the fourth loss calculation to obtain the global loss value.

[0217] The global loss value is used to characterize the overall error of the video understanding model being trained at the overall output level.

[0218] In some embodiments, the fourth loss calculation based on the prediction result and the target output ground truth value to obtain the global loss value can be achieved in the following way: First, according to the preset output format, the target frame description ground truth value, target identifier ground truth value and target segment description ground truth value in the target output ground truth value are summarized into a target output word sequence; the one-hot encodings of the positions of each word in the target output word sequence in the preset dictionary are merged into a third one-hot encoding sequence; the third prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting the output prediction result is obtained; then, the third cross-entropy loss value is determined based on the third prediction probability in the third prediction probability distribution matrix and the third one-hot encoding label in the third one-hot encoding sequence; finally, the third cross-entropy loss value is determined as the global loss value.

[0219] A preset output format refers to pre-defined data organization rules used to define the order, separation method, and output structure of target frame description truth values, target identifier truth values, and target fragment description truth values ​​in a unified sequence. The purpose of the preset output format is to organize various types of supervision data into a single sequence form, facilitating subsequent unified loss calculation. For example, the preset output format can be set to "target frame description truth values ​​first, target identifier truth values ​​in the middle, and target fragment description truth values ​​last," or separator words can be added between each part. The server receives the preset output format, target frame description truth values, target identifier truth values, and target fragment description truth values. Based on the order and separation rules set in the preset output format, it expands the target frame description truth values, target identifier truth values, and target fragment description truth values ​​into continuous words, organizes them uniformly, and outputs the target output word sequence.

[0220] The prediction result is the overall prediction result output by the video understanding model within the current training time window. The prediction result includes the predicted third descriptive text, the predicted second identifier, and the predicted fourth descriptive text. The prediction result uses the same preset output format as the target output word sequence; that is, the predicted third descriptive text, the predicted second identifier, and the predicted fourth descriptive text in the prediction result are also summarized according to the preset output format. The third one-hot encoded sequence is an encoded sequence formed by concatenating the one-hot encoded words in the target output word sequence in the overall output order. The third prediction probability distribution matrix is ​​the set of probability distribution results generated by the video understanding model for all words in the preset dictionary at each output position when generating the prediction result. The third prediction probability is the prediction probability value of any word corresponding to any output position in the third prediction probability distribution matrix. The third one-hot encoded label is the label value used in the third one-hot encoded sequence to indicate the position of the target word. The third cross-entropy loss value is the loss result calculated based on the difference between the third prediction probability and the third one-hot encoded label. The global loss value is used to characterize the overall degree of difference between the prediction result and the target output word sequence.

[0221] In some embodiments, the server takes the prediction result, the target output word sequence, and a preset dictionary as input. First, it converts each word in the target output word sequence into a third one-hot encoded sequence. Then, it reads the third prediction probability distribution matrix of the prediction result at each output position from the video understanding model output. Next, it performs cross-entropy loss calculation based on the third prediction probability in the third prediction probability distribution matrix and the third one-hot encoded label in the third one-hot encoded sequence, outputs the third cross-entropy loss value, and determines the third cross-entropy loss value as the global loss value.

[0222] Each row in the third prediction probability distribution matrix represents the prediction probability distribution of the prediction result at an output position for all words in the preset dictionary. The third one-hot encoded label represents the target word position of the target output word sequence at that output position. The smaller the third cross-entropy loss value, the closer the overall prediction result is to the target output true value. By setting a global loss value, a unified constraint can be applied to the overall output result in addition to keyframe description, video segmentation, and segment description, thereby improving the stability of the video understanding model training process, the consistency between various prediction results, and the accuracy of the final video segmentation task.

[0223] Here, by calculating the global loss value based on the prediction results and the target output word sequence, the overall output of the video understanding model can be uniformly constrained during the training process. This enables the prediction of the third descriptive text, the prediction of the second identifier, and the prediction of the fourth descriptive text to be optimized collaboratively under the same training objective. This reduces the problem of inconsistent overall results despite individual optimization of each output, thereby improving the overall output consistency, format stability, and comprehensive prediction quality of the trained video understanding model.

[0224] Step S4045: Based on the preset first weight, second weight, third weight and fourth weight, the keyframe description loss value, video segmentation loss value, segment description loss value and global loss value are weighted and summed to obtain the multi-task joint loss value.

[0225] The first, second, third, and fourth weights are used to adjust the contribution ratio of different task loss items to the total loss. The keyframe description loss value, video segmentation loss value, segment description loss value, and global loss value are multiplied by their respective weights and then summed in a weighted manner to obtain the multi-task joint loss value.

[0226] Here, the first weight controls the influence of the keyframe description task on the training direction, the second weight controls the influence of the video segmentation task, the third weight controls the influence of the segment description task, and the fourth weight controls the influence of the overall output task on the training direction. Through weighted summation, a balance can be achieved between segmentation accuracy, segment description accuracy, keyframe description accuracy, and overall output consistency. By simultaneously calculating the keyframe description loss, video segmentation loss, segment description loss, and global loss, the video understanding model can be jointly supervised at four levels: keyframe level, segmentation point level, segment level, and overall result level. This allows the video understanding model to simultaneously learn keyframe content understanding, segmentation boundary recognition, segment content summarization, and overall result consistency control during training, thereby improving the semantic understanding ability, segmentation accuracy, description coherence, and result stability of the trained video understanding model in video segmentation tasks.

[0227] By extracting the target frame description truth value for the current keyframe of the sample from the frame description truth value sequence, the target identifier truth value for the cumulative keyframe sequence of the sample from the identifier truth value sequence, and the target segment description truth value for the sample video segment formed after segmentation according to the target true segmentation point from the segment description truth value sequence, the predicted third description text, predicted second identifier, and predicted fourth description text are calculated with real data that are consistent in time range, supervision granularity, and arrangement order, respectively. This reduces the interference of irrelevant sample keyframes, irrelevant true segmentation points, and irrelevant sample video segments on the multi-task joint loss value, improves the accuracy and stability of the multi-task joint loss value, and enhances the consistency between frame-level supervision, segmentation point supervision, and segment-level supervision.

[0228] Step S405: Based on the predicted third description text, the predicted second identifier, and the predicted fourth description text, update the sample cumulative state information to obtain the sample cumulative state information for the next training time window of the current training time window.

[0229] The current training time window and the next training time window partially overlap in the time dimension. After generating the prediction results for the current training time window, instead of directly using the target output ground truth to update the sample accumulated state information, the system updates the sample accumulated state information using the predicted third descriptive text, the predicted second identifier, and the predicted fourth descriptive text. This simulates the process by which the video understanding model under training gradually accumulates context based on its own output during the actual application stage.

[0230] The method for updating the sample cumulative state information can include: writing the predicted third description text into the sample first description text part, writing the predicted second identifier into the sample first identifier part, and writing the predicted fourth description text into the sample second description text part, thereby obtaining the sample cumulative state information for the next training time window. The current training time window and the next training time window partially overlap in the time dimension, so the video understanding model to be trained can still observe sample keyframes near the boundary again in the next training time window, which is beneficial for learning cross-window continuous reasoning and historical correction capabilities. For example, when the current training time window is sample keyframes 7 to 14 and the predicted second identifier is [5, 13], the prediction result is written into the sample cumulative state information, and the updated sample cumulative state information is provided to the next training time window sample keyframes 13 to 20, so that the video understanding model to be trained can continue to perform subsequent reasoning based on the previous round of prediction context during the training phase of sample keyframes 13 to 20.

[0231] Step S406: Based on the multi-task joint loss value, update the network parameters of the video understanding model to be trained to obtain the trained video understanding model.

[0232] The trained video understanding model is used to implement the video segmentation method described above. Backpropagation is performed based on the multi-task joint loss value to obtain the gradient information of each network parameter of the video understanding model. The network parameters of the video understanding model to be trained are then updated according to an optimization algorithm. The optimization algorithm can be the AdamW optimization algorithm or other gradient optimization algorithms. The server sequentially executes steps S401 to S405 on all sample videos in the training sample set according to a preset training time window. After each iteration, the network parameters of the video understanding model are updated based on the multi-task joint loss value until the entire training time window for all sample videos is processed, resulting in the trained video understanding model.

[0233] In some embodiments, the server can also perform multiple rounds of traversal on the training sample set according to a preset training round to further improve the convergence effect and generalization ability of the trained video understanding model. The trained video understanding model is used to implement the above-mentioned video segmentation method, that is, to receive keyframe sequences and accumulated state information during the inference stage, and to output the third descriptive text, the second identifier, and the fourth descriptive text.

[0234] Here, accumulated sample state information and noisy historical context are introduced during the training phase. The video understanding model is constrained by the joint loss of keyframe description, video segmentation, fragment description, and overall output tasks. This enables the model to not only learn keyframe content understanding and segmentation point recognition capabilities, but also to learn continuous reasoning and result correction capabilities when historical descriptions contain errors, historical segmentation has biases, or historical fragment descriptions contain noise. Therefore, the trained video understanding model can better adapt to long video segmentation processing scenarios in practical applications, improving the accuracy of keyframe descriptions, the accuracy of segmentation boundary localization, the coherence of fragment descriptions, and the overall output stability, while enhancing its tolerance and correction capabilities for historical error propagation.

[0235] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.

[0236] This application provides a video plot segmentation method based on cumulative keyframe analysis and correction. This method executes the process sequentially through a complete processing chain: video preprocessing, feature extraction, scene segmentation, keyframe deduplication, plot understanding and segmentation, and post-processing. Specifically, the video preprocessing stage handles video decoding and frame sampling; the feature extraction stage extracts scene embedding features and text in parallel; the scene segmentation stage performs scene boundary detection based on scene features; the keyframe deduplication stage removes visually similar but dialogue-repeating redundant frames; the plot understanding and segmentation stage performs cumulative plot analysis and segmentation based on a large model; and the post-processing stage integrates the segmentation results and plot descriptions to generate the final output.

[0237] In some embodiments, the input to the video preprocessing stage is continuous video, such as an mp4 file, and the output is a uniformly sampled sequence of image frames. The core parameter of the video preprocessing stage is the sampling rate. A sampling rate of 5fps can be used for frame extraction. This parameter can be adjusted, but considering issues such as covering short sentences or long sentences, the sampling rate must be guaranteed to be 5fps. For example, the word "no" spoken rapidly has a duration of 0.3 seconds, requiring the frame to be detected. The selection of the sampling rate is based on the following considerations: a sampling rate that is too low (e.g., 1fps) will lead to the loss of key actions or expressions, while a sampling rate that is too high (e.g., 10fps) will increase the computational burden of subsequent processing. Experiments show that a sampling rate of 5fps can effectively capture key plot points in movies and TV shows while controlling computational costs. This video plot segmentation method can use the VideoCapture object from the OpenCV library to open the video file, obtain the total number of frames and frame rate information, calculate the sampling interval as the frame rate divided by 5, and extract one frame at each sampling interval. For example, for a 25fps video, one frame is extracted every 5 frames; for a 30fps video, one frame is extracted every 6 frames. The extracted frames are saved in JPEG format, and the timestamp information of each frame is recorded for subsequent generation of time location of the segmentation points.

[0238] In some embodiments, the feature extraction stage employs a parallel processing architecture, simultaneously extracting visual and textual features. The visual feature portion uses scene embedding features to characterize the visual content of the scene, including information such as the scene environment, character actions, and composition. A ResNet18 convolutional neural network can be used as the scene classification model. ResNet18 is a lightweight version of a residual network with 18 layers, effectively mitigating the gradient vanishing problem in deep networks through residual connections. ResNet18 is pre-trained on the ImageNet dataset and has excellent visual feature extraction capabilities. Scene classification models for movies and TV shows can also be trained, or a ResNet model trained on the open-source Places365 scene dataset can be directly used. An adaptive modification to ResNet18 is made, removing the fully connected classification layers of the original model and replacing them with scene classification layers, with N classes. The input image is preprocessed (e.g., scaled to 224×224 resolution and normalized), and features are extracted through convolutional layers. After a global average pooling layer, a 512-dimensional feature vector is obtained; this 512-dimensional feature vector is the scene embedding feature. This video segmentation method is a fine-tuned version of ResNet18. It collects about 50,000 typical scene images from movies and TV dramas, including indoor scenes, outdoor scenes, close-up shots, and panoramic shots. It is fine-tuned and trained on the basis of ImageNet pre-trained weights to make ResNet18 more adaptable to the visual style of movies and TV dramas.

[0239] The text features are derived from text extraction. Text extraction employs Optical Character Recognition (OCR) technology to extract the text content of the subtitle area from video frames. Open-source OCR models can be used; OCR models are based on deep learning and have good recognition performance for various characters, including Chinese, English, and numbers. The input to OCR processing is the original video frame, and the output is a text string and the text's position information within the frame. The specific implementation includes text region localization, text content recognition, and text timestamp association. In film and television dramas, the subtitles are generally located in the central area at the bottom of the frame (vertically occupying 1 / 7 of the height) (horizontally, excluding the two outer 1 / 10 edges). A portion of the frame is directly cropped, and then character recognition is performed on the text region to output the text string. The recognized text is then associated with the corresponding timestamp to generate a time-aligned sequence of dialogue. The above text region localization is sufficient for normal television dramas. As an alternative, a dialogue detection model can be trained to replace fixed-region cropping, thereby detecting the position of dialogue in an image. This method is adaptable to various dialogue formats, such as short dramas, vertical short dramas, horizontal screens, and square screens. Visual and textual features are extracted in parallel at this stage, providing complementary information for subsequent scene segmentation, duplicate frame recognition, character naming correction, and plot understanding. In terms of visual segmentation, the scene segmentation stage performs scene boundary detection based on scene embedding features, dividing a continuous video frame sequence into multiple scenes. Each scene contains a set of visually similar and temporally continuous frames. For two adjacent frames i and i+1, the cosine similarity between the scene embedding vector emb_i of frame i and the scene embedding vector emb_{i+1} of frame i+1 is calculated, using the formula similarity(i, i+1) = (emb_i•emb_{i+1}) / (||emb_i|| × ||emb_{i+1}||). Here, • represents the vector dot product, and || represents the L2 norm of the vector. The cosine similarity ranges from [-1, 1], with a higher value indicating greater similarity between the two frames.

[0240] The scene boundary determination uses a normalized adaptive thresholding method, with different similarity thresholds for different TV series. The specific algorithm is as follows: First, calculate the cosine similarity of all adjacent frame pairs to obtain a cosine similarity sequence. Then, collect the similarity s0 of a small number of consecutive frames. For example, collect the cosine similarity of the first episode of each of 50 TV series with evenly distributed themes, and calculate the mean μ and standard deviation σ of the similarity sequence for each series. Then, normalize the similarity sequence for each series using the formula s = (s0 - μ) / σ. Next, extract the cosine similarity of consecutive frames of approximately 20 transition shots from each of the 50 TV series (the scene label under this cosine similarity is 1, indicating different scenes). Simultaneously, extract the cosine similarity of consecutive frames within the same scene (the scene label under this cosine similarity is 0, indicating the same scene), forming a training set and training an SVM support vector machine classifier to obtain the optimal classification threshold thr for consecutive frames.

[0241] In other words, for a given series, all keyframes from the first episode can be extracted (keyframes are those generated by the extraction method described above; for efficiency, frames representing half the video length can also be used). Scene embedding vectors for these keyframes are extracted, and the mean μ and standard deviation σ of the cosine similarity between the scene embedding vectors of consecutive frames are calculated. For each subsequent keyframe in the series, scene embedding vectors are extracted, the cosine similarity between consecutive frames is calculated, and the cosine similarity is normalized. Then, a similarity threshold thr is used to determine whether it represents a scene boundary. This video segmentation method can adapt to the visual style differences of various videos. For action films with drastic visual changes, the similarity threshold thr automatically decreases; for art films with gentle visual changes, the similarity threshold thr automatically increases.

[0242] Scene boundary determination can also be achieved using open-source scene segmentation tools, eliminating the need for the aforementioned scene classification model to extract scene embedding features. This tool offers various segmentation algorithms, such as content-difference-based and threshold-based algorithms. For example, when using a content-difference-based segmentation algorithm, scene segmentation detection can be performed based on histogram differences in the HSV color space, where the segmentation threshold can be set to 30.0 and the minimum scene length can be set to 15 frames.

[0243] After scene segmentation, keyframe deduplication is performed. The resulting frame sequence contains visually similar but dialogue-repeating redundant frames, a common occurrence in films and television dramas, such as repeated shots and flashbacks. Keyframe deduplication aims to remove these redundant frames and retain representative keyframes. Duplicate frame detection is based on two conditions: visual similarity and dialogue similarity. The visual similarity condition is that the cosine similarity between the scene embedding features of two frames is greater than a threshold θ (vis), for example, 0.9. The dialogue similarity condition is that the OCR text content of two frames is completely identical, or the normalized edit distance is less than a threshold θ (for example, for a 10-character line with two different characters, the normalized similarity is 8 / 10). Only frames meeting either of these two conditions are considered duplicates. For detected duplicate frame groups, the first frame in the group is prioritized for retention, as it contains the most complete contextual information. If the OCR recognition quality of the first frame is lower than that of the second frame (e.g., the text is too short, contains noise, or cannot recognize characters), then the frame with the longest OCR recognition quality within the duplicate frame group is retained. The deduplicated keyframe sequence is denoted as {F_1, F_2, ..., F_N}, where N is the number of keyframes. Each keyframe contains a frame identifier, timestamp, scene embedding features, OCR text, and the original image.

[0244] For the plot understanding and segmentation stage, an open-source large model can be used to perform semantic understanding and plot segmentation on keyframe sequences. This model can accurately understand elements such as characters, actions, and scenes in keyframes and generate coherent text descriptions based on context. To better adapt the large model to the video plot segmentation task, supervised fine-tuning (SFT) can be performed. Note: An untrained large model cannot correctly complete this video plot segmentation task because it requires handling complex editing situations and deep semantic understanding. Relying solely on cue words is insufficient to accurately represent the video plot segmentation task and yields poor results.

[0245] To this end, 2000 representative video segmentation samples can be collected, along with several hundred complex samples containing complex scenes and various editing techniques. Each sample includes keyframes of varying lengths of original video scenes (e.g., 1 to 10 scenes, each scene containing 1 to 10 frames), reference character information (e.g., names and images of main characters), keyframe sequences (e.g., representative frames after 3fps sampling), manually annotated cumulative plot segmentation points (frame identifier list), manually annotated cumulative plot descriptions (textual descriptions of each segment), keyframe descriptions of historical scenes, historical plot segmentation points (with random error examples constructed to randomly modify segmentation point positions), and historical cumulative plot descriptions (textual descriptions of each segment, with a 0.5 probability of constructing one random name error for each scene). The samples can cover various video types, including period dramas, modern dramas, suspense dramas, and comedies, ensuring that the large model can learn the narrative patterns of different styles of film and television dramas.

[0246] To ensure consistency between the training and inference processes, a prompt word template was designed. For example, the prompt word template could be: Given a TV series or movie, a sequence of keyframes for a certain part of the plot is provided in chronological order. Descriptive text for historical keyframes and character reference images are also given. Referring to the descriptive text of the historical keyframes, the historical plot segmentation points, and the descriptive text of the historical accumulated plot, the prompt word should provide the descriptive text of the current keyframe (i.e., the current description), the current accumulated plot segmentation point, and the descriptive text of the current keyframe after corrections such as dialogue (i.e., the correct current description). This should be returned in the following JSON format: {'Current Description':[x,x,x,x,x,x,x,x,x,x,x], 'Current Accumulated Plot Segmentation Point':[x,x,x], 'Correct Current Description':[x,x,x,x,x,x,x,x,x,x,x], 'Descriptive Text of Current Accumulated Plot':[x,x,x,x]}. Each element in the current accumulated plot segmentation point is the frame identifier of the segmented frame. For example, assuming the current cumulative plot breakpoint is [7,14], it means the breakpoint starts at frames 7 and 14. That is, frames 1 to 6 constitute one plot, frames 7 to 13 constitute another, and frames 14 to the last frame constitute yet another. If the current content is the same plot, the returned list of current cumulative plot breakpoints can be empty. Each item in the current description is the descriptive text for the keyframes of the storyline from images 1 to 10. The descriptive text must include the characters, environment, and events, and requires careful comparison of the characters in the character reference image and the storyline images to ensure accuracy. The descriptive text for the current cumulative plot provides a corresponding number of text descriptions based on the number of plots formed by the cumulative plot breakpoints. For characters addressed by name in the keyframes, the large model must directly reference the character name. For characters whose names are directly mentioned in the keyframes, the character name should also be used directly. For uncertain or unfamiliar people, use "classmate a," "passerby a," "passerby b," etc., as substitutes. The cue words further require distinguishing between dialogue and advertising text in keyframes. Generally, advertisements, show titles, and platform names are located in the top left, bottom left, top right, and bottom right corners of the screen. This information does not need to appear in the current description or plot text. The cue words also require the large model to analyze the plot after outputting the current keyframe content, reflect on the current content, and correct character identities and dialogue. If a character with a clear title appears in the context, such as being addressed as A in dialogue, but the large model analyzes their identity as B, the large model needs to correct its previous identity analysis, using the character's identity as defined in the current keyframe's dialogue. If the context characters are consistently A and C, and suddenly character B appears in a frame, the large model needs to carefully check whether it incorrectly identified another character as B, or whether a necessary plot segmentation was not performed in a frame before character B appeared, and then correct accordingly. After verification, the result in the correct current description must be the correct result.In the initial input scenario, there is no descriptive text for historical keyframes, no historical plot segmentation points, and no descriptive text for historical cumulative plots. This design unifies character recognition, dialogue understanding, segmentation decisions, and self-correction into a single output structure, and also ensures strict alignment between SFT data construction and online inference formats.

[0247] For fine-tuning, Low-Rank Adaptation (LoRA) can be used for efficient parameter fine-tuning, or full-parameter fine-tuning can be performed, that is, fine-tuning all parameters of a large model. When using LoRA fine-tuning, LoRA adds a low-rank matrix to the linear layers of the original large model and only trains the newly added parameters in the low-rank matrix, thereby significantly reducing the computational cost of fine-tuning. The LoRA configuration can be set as follows: rank r = 64, target modules are the query projection module, key projection module, value projection module, and output projection module in the attention mechanism, learning rate = 2 × 10^-4, batch size = 8, training epochs = 3, and the optimizer is AdamW. The loss function adopts multi-task learning loss, and the output of each task includes four parts: plot segmentation, plot description, keyframe description, and all answers. The loss consists of four components: plot segmentation loss L_cut, which classifies plot segmentation points using binary cross-entropy loss; plot description loss L_plot, which calculates the language model loss for the generated plot descriptions; keyframe description loss L_frame, which calculates the language model loss for the description of each keyframe; and all-answer loss L_0, which constrains the output format and the three output components: plot segmentation, plot description, and keyframe description. The total loss can be expressed as: L_total = λ_1•L_cut + λ_2•L_plot + λ_3•L_frame + λ_4•L_0, where λ_1 = 0.4, λ_2 = 0.3, λ_3 = 0.1, and λ_4 = 0.2, used to balance the learning of the four tasks, ensuring correct plot segmentation and plot description, while allowing for incomplete keyframe descriptions.

[0248] When using full-parameter fine-tuning, each SFT training data point is stored in a uniform format, including a data sequence number, task prompt words, and a standard answer. The training process can be configured with the following parameters: learning rate of 5e-5, batch size of 8, gradient accumulation steps of 4, global batch size of 32, AdamW optimizer, and 2 training epochs. Model parameters are initialized using pre-trained open-source model network weights. During training, the text content from the task prompt words is used as model input. The large model predicts the output text word by word, obtaining a prediction probability matrix. This prediction probability matrix is ​​then compared with the standard answer text in the samples to calculate the supervised fine-tuning loss. The supervised fine-tuning loss uses cross-entropy loss. The one-hot vector of the correct word in the supervised text in the dictionary is used as the supervised signal; where the label of the i-th sample at the j-th position in the answer within the batch can be represented as The prediction probability of each word by the large model is denoted as... Cross-entropy loss is used to improve the model's prediction probability of the target word. This indicates the number of samples in each batch. This represents the number of predicted words for the i-th sample within the batch. Both low-rank adaptation fine-tuning and full-parameter fine-tuning can be retained simultaneously, allowing for flexible selection of the appropriate training scheme based on computing resources, target model size, and iteration cycle.

[0249] Since the number of keyframes exceeds the context length limit of large models, it is difficult to complete inference by inputting all video frames at once. Therefore, a cumulative inference strategy can be adopted when applying the model, processing the keyframe sequence frame by frame or batch by batch. The initial state is set as follows: the description text of the historical accumulated plot is an empty list, the historical plot segmentation points are an empty list, and the current description is an empty string.

[0250] Subsequently, for the i-th keyframe (i from 1 to N), the following operations are performed: Construct a cue word template by filling it with the character list, descriptions of historical keyframes, historical plot breakpoints, and descriptions of historical cumulative plots. Obtain the current batch of keyframe sequences, with a maximum of 10 frames input to avoid exceeding the image processing capabilities of the large model. Character reference images can also be added. The filled cue words can be input into the large model to obtain JSON format output, which includes the current description, the current cumulative plot breakpoint, the correct current description, and the description text of the current cumulative plot. After completing one inference iteration, add the correct current description to the description text of historical keyframes, update the historical cumulative plot breakpoints using the current cumulative plot breakpoints, and update the description text of historical cumulative plots using the description text of the current cumulative plot. If the current cumulative plot breakpoint in the output is not empty, replace the historical cumulative plot breakpoint with the new current cumulative plot breakpoint. If a new description text of the current cumulative plot is output, replace the list of historical cumulative plot description texts with the new current cumulative plot description text. For continuous video processing, a sliding time window strategy is used. The time window size can be set to 10 frames, with a step size of 2 frames, meaning adjacent time windows overlap by 2 frames. If there is more video memory, the window size can be further increased; for example, with a single card performing 10 frames of inference, a 4-card system supporting data-parallel inference can expand the window to 40 frames. Sliding time window processing ensures plot continuity while reducing the computational burden of a single inference iteration.

[0251] The model possesses self-correcting capabilities during the application phase, primarily manifested in character identity correction and plot segmentation correction. Regarding character identity correction, during frame-by-frame processing, the large model may misidentify certain characters, such as misidentifying protagonist A as supporting character B. When a clear character name appears in subsequent frames, such as mentioning Xiao Wang in dialogue, the large model can detect the inconsistency through a reflection mechanism and correct the previous misidentification. The prompts explicitly require the large model to correct character identities, checking if the character's name and identity are consistent with the context. If inconsistent, the latest dialogue information takes precedence, and the character names in the historical descriptions are updated simultaneously. Regarding plot segmentation correction, plot segmentation is a cumulative decision-making process; early incorrect segmentation decisions may be made due to incomplete information. As the plot unfolds, the large model acquires more contextual information and can identify previous unreasonable segmentations. For example, at frame 10, the large model might suggest plot segmentation, but at frame 15, it might discover that frames 10 through 15 actually belong to the same plot. In this case, the video plot segmentation method can remove the plot segmentation point from frame 10 during cumulative plot segmentation, thus maintaining plot integrity. The output at each inference iteration is the cumulative plot segmentation point, not the segmentation point of the current keyframe sequence, allowing for adjustments to all previous segmentation decisions at any time. This mechanism gives the video plot segmentation method strong fault tolerance and enhances the global narrative coherence in continuous video processing.

[0252] To further improve usability during the application phase, fault-tolerant designs were implemented to address common issues such as text interference, repeated shots, and flashbacks. Regarding text interference, frames contain interfering text such as advertising slogans, platform watermarks, and program logos. Related technologies require training specialized text cleaning models to remove this interference. This video plot segmentation method uses prompts to allow a large model to identify the location features of this text (i.e., its position in the four corners of the frame) and explicitly instructs the large model to ignore this information. This rule-based instruction is more efficient and flexible than training a specialized model. Regarding repeated shots and flashbacks, TV series and movies use narrative techniques such as flashbacks and repeated shots, which related technologies sometimes misclassify as new plot points. This video plot segmentation method, through keyframe deduplication combined with contextual inference in plot understanding, can identify these repeated contents and correctly categorize them into the original plot, avoiding incorrect plot segmentation. For character-specific plot preview services, the video plot segmentation method can also combine character reference images and character lists to filter out plot segments containing the male lead, female lead, or other specified characters during post-processing, thereby generating character-specific plot preview results.

[0253] After all inference is complete, the video plot segmentation method can continue to perform global consistency checks, timestamp mapping, result formatting, and plot summary generation. By default, the confidence level of each global segmentation point is 1. Then, it checks whether the time interval between plot segmentation points is reasonable. For example, if the interval between adjacent segmentation points is too short (i.e., the duration of a plot is less than 2 seconds), the segmentation confidence level before and after the corresponding segmentation point is marked as 0.5. In practical applications, such as plot previews and plot markers, positions with a confidence level below 1 can be left unmarked and not marked for plot previews. It can also check the coherence of plot descriptions. For example, if character A and character B appear in multiple frames before and after a scene, and then character C suddenly appears in one frame, the plot confidence level of the corresponding segment is set to 0.3. Plot confidence levels below 1 can be excluded from output to plot preview applications.

[0254] Subsequently, based on the frame identifier and timestamp mapping relationship recorded during the video preprocessing stage, the frame identifier-level segmentation points can be converted into time positions. The calculation formula is: timestamp = frame identifier / sampling rate. For example, a segmentation point with a frame identifier of 150 corresponds to a timestamp of 150 / 5 = 30 seconds at 5fps sampling.

[0255] Meanwhile, the plot description can be further compressed to obtain a more suitable summary for browsing. The plot description output by the plot understanding tool is relatively detailed, and a concise summary can be generated in the post-processing stage. An extraction-based summarization method is used to extract key information from the plot description, including main characters, key actions, and highlights, generating a one-sentence plot summary. For quick-viewing scenarios, this one-sentence plot summary (501) can be directly displayed to the user, such as... Figure 12 As shown. For operational tagging scenarios, a one-sentence plot summary can be directly used as candidate text. For character quick preview scenarios, a one-sentence plot summary can also be displayed together with character names to form a catalog-style result that can be quickly browsed.

[0256] Based on the aforementioned processing chain, this video plot segmentation method offers several beneficial effects: it achieves accurate plot segmentation, surpassing scene segmentation. Through reliable keyframe extraction and large-scale model accumulation and understanding, it segments plots based on semantic understanding of keyframes, more accurately identifying plot boundaries rather than just physical scene transitions, outperforming methods based on visual features. It possesses self-correcting capabilities and strong fault tolerance. By employing a cumulative inference strategy, the large model can correct previous judgments based on newly acquired information, resolving issues such as character recognition errors and historical segmentation position errors, while maintaining stronger plot coherence. Furthermore, it reduces reliance on precise preprocessing models. Related technologies require strong transition frame recognition and audio understanding capabilities, or rely on basic video understanding models, such as character recognition models and text tag removal models, for preprocessing. This video plot segmentation method, through semantic understanding of the large model and cue word engineering, tolerates preprocessing errors to a certain extent, thereby reducing the need for precise models and lowering deployment costs. It also exhibits strong scalability, adapting to various video types. Both the normalized scene segmentation method and the segmentation method based on advanced semantic plot understanding do not rely on prior knowledge of specific video types. Through fine-tuning, they can adapt to different styles and types of video content. Overall, the video plot segmentation method can provide stable plot-level segmentation capabilities across various video content types, including TV dramas, online short dramas, variety show clips, and movie clips. It can also directly serve multiple business scenarios such as plot previews, character plot previews, plot point analysis, content understanding, and video summarization.

[0257] like Figure 13 As shown, this video plot segmentation method can be directly applied to video playback scenarios such as plot previews (including plot summaries 601) and plot previews of specific characters, and it achieves better segmentation results compared to related technologies. (For comparison...) Figure 1 Taking the relevant technologies in the video as an example, the segmentation point 701 obtained by this video plot segmentation method (such as...) Figure 14 (As shown) is frame 5, meaning frames 1 to 4 are correctly merged into a transition sequence. Figure 1The relevant technology incorrectly splits frames 1 to 4 into three scenes (i.e., storyboards), with frames 5 to 7 representing a subsequent scene. The output information of this video scene segmentation method includes corrected frame descriptions, scene segmentation analysis, and scene segmentation points. The corrected frame descriptions are as follows: Frame 1 shows a modern skyscraper under a clear, sunny sky; the scene is a cityscape. Frame 2 switches to a busy highway with vehicles in motion; the scene changes from a cityscape to a traffic scene. Frame 3 switches to an overhead view of an overpass or railway bridge, with sunlight filtering through the bridge structure; the scene changes from a traffic scene to a building structure scene. Frame 4 switches to a panoramic view of the city, showing skyscrapers and green vegetation; the scene changes from building structures to a city panorama. Frame 5 suddenly switches to a forest or woods where a character appears, seemingly walking or observing; the scene changes from a city panorama to a natural environment. Frame 6 focuses on the side of the character, who seems to be thinking or observing; the scene continues in the woods. In frame 7, the character begins to speak, and the subtitles indicate that it was sometime three years ago. The scene is still a forest, but dialogue has been added. The plot segmentation analysis shows that frames 1 to 4 display different scenes of the city, including skyscrapers, highways, overpasses, and city panoramas, which belong to the same plot, namely, the display of the urban environment. Starting from frame 5, the scene suddenly switches to a forest, characters appear, and there is dialogue. This is obviously a new plot, and the corresponding plot segmentation point is [5], such as Figure 14 As shown in the figure, this result demonstrates that the proposed video plot segmentation method can not only handle obvious scene changes but also identify the narrative continuity within transition segments, thus providing more stable and narrative-consistent point-based results for plot quick previews and character plot quick previews.

[0258] It is understood that in the embodiments of this application, if data related to user information or enterprise information is involved, when the embodiments of this application are applied to specific products or technologies, it is necessary to obtain user permission or consent, or to obfuscate this information in order to eliminate the correspondence between this information and the user; and the collection and processing of related data should strictly comply with the requirements of relevant laws and regulations when applied in practice, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0259] The following continues to describe the exemplary structure of the video segmentation device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 3As shown, the video segmentation device 455 includes: a first iterative processing module 4551, used to perform the following iterative processing on the keyframes of the target video according to a preset time window: obtaining the current keyframe sequence and accumulated state information of the target video within the current time window; the accumulated state information includes: a first description text for each historical keyframe in the historical keyframe sequence before the current time window, a first identifier of the historical segmentation point corresponding to the historical keyframe sequence, and a second description text for each video segment obtained after segmenting the historical keyframe sequence according to the historical segmentation point; and determining a third description text for each current keyframe in the current keyframe sequence based on the current keyframe sequence, the first description text, the first identifier, and the second description text. The video segmentation module 4552 is used to perform video segmentation on the target video based on the cumulative segmentation information obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation information, the second identifier of the cumulative segmentation point corresponding to the cumulative keyframe sequence, and the fourth description text of each video segment obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation point; the cumulative keyframe sequence includes the current keyframe sequence and the historical keyframe sequence; based on the third description text, the second identifier and the fourth description text, the cumulative state information is updated to obtain the cumulative state information of the next time window of the current time window; the current time window and the next time window partially overlap in the time dimension; the video segmentation module 4552 is used to perform video segmentation on the target video based on the cumulative state information obtained after the iterative processing of all keyframes in the target video, and obtain the video segmentation result of the target video.

[0260] In some embodiments, the first iterative processing module 4551 is further configured to: obtain object reference information corresponding to the target video, and text information corresponding to each current keyframe in the current keyframe sequence; fill the object reference information, text information, first descriptive text, first identifier, and second descriptive text into a preset prompt word template to obtain the target prompt word; input the current keyframe sequence and the target prompt word into a pre-trained video understanding model; and determine the third descriptive text, the second identifier, and the fourth descriptive text through the video understanding model.

[0261] In some embodiments, the first iterative processing module 4551 is further configured to: generate initial results for the current keyframe sequence using a video understanding model; the initial results include initial frame description text for each current keyframe in the current keyframe sequence, initial identifiers of initial cumulative segmentation points corresponding to the cumulative keyframe sequence, and initial segment description text for each first video segment obtained after segmenting the cumulative keyframe sequence according to the initial cumulative segmentation points; calibrate the initial frame description text based on preset calibration text in the target prompt words to obtain a third description text; calibrate the initial identifier to obtain a second identifier; and calibrate the initial segment description text based on the second identifier to obtain a fourth description text.

[0262] In some embodiments, the first iterative processing module 4551 is further configured to: for each current keyframe in the current keyframe sequence, obtain the current text corresponding to the current keyframe from the target prompt words; when the current text includes a first object identifier of the target object in the current keyframe, determine a second object identifier of the target object in the current keyframe from the initial frame description text; when the first object identifier is different from the second object identifier, replace the second object identifier in the initial frame description text with the first object identifier to obtain a third description text; when the current text does not include any first object identifier, or when the first object identifier is the same as the second object identifier, determine the initial frame description text as the third description text.

[0263] In some embodiments, the first iterative processing module 4551 is further configured to: for any initial cumulative segmentation point, determine the semantic coherence between the preceding and following video segments corresponding to the initial cumulative segmentation point based on the third descriptive text and the first descriptive text in the target prompt word; when the semantic coherence indicates semantic coherence between the preceding and following video segments, delete the initial identifier of the initial cumulative segmentation point; when the semantic coherence indicates semantic incoherence between the preceding and following video segments, determine the initial identifier of the initial cumulative segmentation point as the second identifier.

[0264] In some embodiments, the first iterative processing module 4551 is further configured to: segment the accumulated keyframe sequence according to the initial accumulated segmentation point corresponding to the second identifier to obtain a plurality of second video segments; for each second video segment, when the second video segment includes a plurality of first video segments, semantically merge the initial segment description texts of the plurality of first video segments to obtain a fourth description text of the second video segment; when the second video segment includes a first video segment, determine the initial segment description text of the first video segment as the fourth description text of the second video segment.

[0265] In some embodiments, the first iterative processing module 4551 is further configured to: concatenate the third description text after the first description text in the cumulative state information to obtain a merged description text; determine the merged description text as the first description text in the cumulative state information of the next time window; determine the second identifier as the first identifier in the cumulative state information of the next time window; and determine the fourth description text as the second description text in the cumulative state information of the next time window.

[0266] In some embodiments, the video segmentation module 4552 is further configured to: obtain an identifier sequence corresponding to the identifiers of all segmentation points of the target video from the accumulated state information obtained after iterative processing; perform video segmentation on the target video sequentially based on the identifiers in the identifier sequence to obtain multiple target video segments; and determine the multiple target video segments as the video segmentation result of the target video.

[0267] In some embodiments, the video segmentation device 455 further includes a keyframe acquisition module; the keyframe acquisition module is configured to: extract frames from the target video according to a preset sampling rate to obtain an initial video frame sequence; determine the scene embedding features and text information of each initial video frame in the initial video frame sequence; segment the initial video frame sequence into multiple video segments based on the scene embedding features; and perform deduplication processing on the initial video frames in each video segment based on the scene embedding features and text information to obtain the keyframe sequence of the target video.

[0268] In some embodiments, the keyframe acquisition module is further configured to: determine the feature similarity between the scene embedding features of each of the two adjacent initial video frames in the initial video frame sequence; normalize the feature similarity to obtain a normalized feature similarity; when the normalized feature similarity is less than a preset first feature similarity threshold, determine the later initial video frame among the two adjacent initial video frames as the segmentation point; and segment the initial video frame sequence based on the segmentation point to obtain multiple video segments.

[0269] In some embodiments, the keyframe acquisition module is further configured to: extract a reference video frame sequence of a preset duration from the target video; determine the reference feature similarity between the scene embedding features of each two adjacent reference video frames in the reference video frame sequence; determine the mean similarity and standard deviation of the similarity of multiple reference feature similarities; for each feature similarity of each two adjacent initial video frames in the initial video frame sequence, determine the difference between the feature similarity and the mean similarity; and determine the ratio of the difference to the standard deviation of similarity as the normalized feature similarity.

[0270] In some embodiments, the keyframe acquisition module is further configured to: for each video segment, when any two adjacent initial video frames in the video segment meet at least one of the following preset conditions, determine the text quality of the text information of each of the two adjacent initial video frames; wherein, the preset conditions are: the feature similarity between the scene embedding features of each of the two adjacent initial video frames in the video segment is greater than a preset second feature similarity threshold; the text similarity between the text information of each of the two adjacent initial video frames is greater than a preset text similarity threshold; when the text quality corresponding to the earlier initial video frame among any two adjacent initial video frames is greater than or equal to a preset quality threshold, retain the earlier initial video frame; when the text quality corresponding to the earlier initial video frame is less than the quality threshold, retain the initial video frame corresponding to the largest text quality among any two adjacent initial video frames; and sort the retained initial video frames in chronological order to obtain the keyframe sequence of the target video.

[0271] The following continues to describe the exemplary structure of the training device 456 for the video understanding model provided in the embodiments of this application, implemented as a software module. In some embodiments, such as... Figure 4 As shown, the training device 456 for the video understanding model includes: a second iterative processing module 4561, used to perform the following iterative processing on the sample keyframes of the sample videos in the training sample set according to a preset training time window: obtaining the current keyframe sequence of the sample videos within the current training time window, sample cumulative state information, and target output ground truth; the sample cumulative state information includes: a sample first description text for each sample historical keyframe in the sample historical keyframe sequence located before the current training time window, a sample first identifier for the sample historical segmentation point corresponding to the sample historical keyframe sequence, and the result obtained after segmentation according to the sample first identifier. The target output truth value includes: a second description text for each sample video segment; the target output truth value includes: a frame description truth value sequence for the sample keyframes of the sample video, a truth value sequence for the identifiers of the actual segmentation points of the sample video, and a segment description truth value sequence for the sample video segments obtained after segmenting the sample video according to the actual segmentation points; noise injection processing is performed on the first description text of the sample to obtain a noisy first description text; noise injection processing is performed on the first identifier of the sample to obtain a noisy first identifier; noise injection processing is performed on the second description text of the sample to obtain a noisy second description text; based on the current keyframe sequence of the sample, the noisy... The first descriptive text, the noisy first identifier, and the noisy second descriptive text are used by a video understanding model to determine a prediction result. The prediction result includes: a predicted third descriptive text for each current keyframe in the current keyframe sequence of the sample; a predicted second identifier for the cumulative segmentation point corresponding to the cumulative keyframe sequence of the sample; and a predicted fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence of the sample according to the cumulative segmentation point. The cumulative keyframe sequence of the sample includes the current keyframe sequence of the sample and the historical keyframe sequence of the sample. Based on the predicted third descriptive text, the predicted second... The cumulative sample state information is updated using the identifier and the predicted fourth descriptive text to obtain the cumulative sample state information for the next training time window of the current training time window; the current training time window and the next training time window partially overlap in the time dimension; loss is calculated based on the predicted third descriptive text and the frame description ground truth sequence, the predicted second identifier and the identifier ground truth sequence, and the predicted fourth descriptive text and the segment description ground truth sequence to obtain a multi-task joint loss value; the network parameters of the video understanding model to be trained are updated based on the multi-task joint loss value to obtain the trained video understanding model.

[0272] In some embodiments, the second iterative processing module 4561 is further configured to: obtain the target frame description truth value for each current keyframe of the sample from the frame description truth value sequence; obtain the target identifier truth value of the target true segmentation point corresponding to the sample cumulative keyframe sequence from the identifier truth value sequence; obtain the target segment description truth value of each sample video segment obtained by segmenting the sample cumulative keyframe sequence according to the target true segmentation point from the segment description truth value sequence; and perform loss calculation based on the predicted third description text and the target frame description truth value, the predicted second identifier and the target identifier truth value, and the predicted fourth description text and the target segment description truth value to obtain the multi-task joint loss value.

[0273] In some embodiments, the second iterative processing module 4561 is further configured to: perform a first loss calculation based on the predicted third descriptive text and the target frame descriptive truth value to obtain a keyframe descriptive loss value; perform a second loss calculation based on the predicted second identifier and the target identifier truth value to obtain a video segmentation loss value; perform a third loss calculation based on the predicted fourth descriptive text and the target segment descriptive truth value to obtain a segment descriptive loss value; perform a fourth loss calculation based on the prediction result and the target output truth value to obtain a global loss value; and perform a weighted summation of the keyframe descriptive loss value, the video segmentation loss value, the segment descriptive loss value, and the global loss value based on preset first weights, second weights, third weights, and fourth weights to obtain a multi-task joint loss value.

[0274] In some embodiments, the second iterative processing module 4561 is further configured to: merge the one-hot encodings of the positions of each word in the target frame description ground truth in the preset dictionary into a first one-hot encoding sequence; obtain the first prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting the third description text; determine the first cross-entropy loss value based on the first prediction probability in the first prediction probability distribution matrix and the first one-hot encoding label in the first one-hot encoding sequence; and determine the first cross-entropy loss value as the keyframe description loss value.

[0275] In some embodiments, the second iterative processing module 4561 is further configured to: calculate the cross-entropy loss based on the predicted second identifier and the true value of the target identifier for each sample cumulative segmentation point corresponding to the sample cumulative keyframe sequence, and obtain the cross-entropy loss value; and aggregate the cross-entropy loss values ​​of all sample cumulative segmentation points corresponding to the sample cumulative keyframe sequence to obtain the video segmentation loss value.

[0276] In some embodiments, the second iterative processing module 4561 is further configured to: merge the one-hot codes of the positions of each word in the target fragment description ground value in the preset dictionary into a second one-hot coding sequence; obtain the second prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting the fourth descriptive text; determine the second cross-entropy loss value based on the second prediction probability in the second prediction probability distribution matrix and the second one-hot coding label in the second one-hot coding sequence; and determine the second cross-entropy loss value as the fragment description loss value.

[0277] In some embodiments, the second iterative processing module 4561 is further configured to: summarize the target frame description truth value, target identifier truth value, and target segment description truth value in the target output truth value into a target output word sequence according to a preset output format; merge the one-hot encodings of the positions of each word in the target output word sequence in the preset dictionary into a third one-hot encoding sequence; obtain the third prediction probability distribution matrix generated by the video understanding model for each word in the preset dictionary when predicting the prediction result; determine the third cross-entropy loss value based on the third prediction probability in the third prediction probability distribution matrix and the third one-hot encoding label in the third one-hot encoding sequence; and determine the third cross-entropy loss value as the global loss value.

[0278] It should be noted that the description of the apparatus in this application embodiment is similar to the description of the method embodiment described above, and has similar beneficial effects as the method embodiment; therefore, it will not be repeated. For technical details not disclosed in this apparatus embodiment, please refer to the description of the method embodiment of this application for understanding.

[0279] This application provides an electronic device, including: a memory for storing computer-executable instructions or computer programs; and a processor for executing the computer-executable instructions or computer programs stored in the memory to implement the above-described method.

[0280] This application provides a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium; wherein, when the processor of an electronic device reads the computer-executable instructions or the computer program from the computer-readable storage medium and executes the computer-executable instructions or the computer program, the above-described method is implemented.

[0281] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs, which are used to cause a processor to execute the computer-executable instructions or computer programs to implement the above-described method.

[0282] In some embodiments, the storage medium may be a computer-readable storage medium, such as ferromagnetic random access memory (FRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or compact disk-read-only memory (CD-ROM); or it may be a device that includes one or any combination of the above-mentioned memories.

[0283] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0284] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file containing other programs or data, for example, in one or more scripts within a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code sections). As an example, executable instructions may be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0285] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A video segmentation method, characterized in that, The method includes: According to a preset time window, perform the following iterative processing on the keyframes of the target video: Obtain the current keyframe sequence and cumulative status information of the target video within the current time window; the cumulative status information includes: a first description text for each historical keyframe in the historical keyframe sequence before the current time window, a first identifier of the historical segmentation point corresponding to the historical keyframe sequence, and a second description text for each video segment obtained after segmenting the historical keyframe sequence according to the historical segmentation point; Based on the current keyframe sequence, the first description text, the first identifier, and the second description text, a third description text for each current keyframe in the current keyframe sequence, a second identifier for the cumulative segmentation point corresponding to the cumulative keyframe sequence, and a fourth description text for each video segment obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation point; the cumulative keyframe sequence includes the current keyframe sequence and the historical keyframe sequence. Based on the third description text, the second identifier, and the fourth description text, the cumulative state information is updated to obtain the cumulative state information of the next time window of the current time window; the current time window and the next time window partially overlap in the time dimension. When the iterative processing of all keyframes in the target video is completed, the target video is segmented based on the accumulated state information obtained after the iterative processing to obtain the video segmentation result of the target video.

2. The method according to claim 1, characterized in that, The step of determining, based on the current keyframe sequence, the first descriptive text, the first identifier, and the second descriptive text, a third descriptive text for each current keyframe in the current keyframe sequence, a second identifier for the cumulative segmentation point corresponding to the cumulative keyframe sequence, and a fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation point, includes: Obtain object reference information corresponding to the target video, and text information corresponding to each current keyframe in the current keyframe sequence; The object reference information, the text information, the first description text, the first identifier, and the second description text are filled into a preset prompt word template to obtain the target prompt word; The current keyframe sequence and the target cue words are input into a pre-trained video understanding model; The video understanding model is used to determine the third descriptive text, the second identifier, and the fourth descriptive text.

3. The method according to claim 2, characterized in that, The step of determining the third descriptive text, the second identifier, and the fourth descriptive text through the video understanding model includes: The video understanding model generates an initial result for the current keyframe sequence. The initial result includes an initial frame description text for each current keyframe in the current keyframe sequence, an initial identifier for the initial cumulative segmentation point corresponding to the cumulative keyframe sequence, and an initial segment description text for each first video segment obtained after segmenting the cumulative keyframe sequence according to the initial cumulative segmentation point. Based on the preset calibration text in the target prompt, the initial frame description text is calibrated to obtain the third description text; the initial identifier is calibrated to obtain the second identifier; The initial fragment description text is calibrated based on the second identifier to obtain the fourth description text.

4. The method according to claim 3, characterized in that, The calibration of the initial frame description text to obtain the third description text includes: For each current keyframe in the current keyframe sequence, the current text corresponding to the current keyframe is obtained from the target prompt word; When the current text includes a first object identifier of the target object in the current keyframe, the second object identifier of the target object in the current keyframe is determined from the initial frame description text; When the first object identifier is different from the second object identifier, the second object identifier in the initial frame description text is replaced with the first object identifier to obtain the third description text; When the current text does not include any first object identifier, or when the first object identifier is the same as the second object identifier, the initial frame description text is determined as the third description text.

5. The method according to claim 4, characterized in that, The calibration of the initial identifier to obtain the second identifier includes: For any initial cumulative segmentation point, based on the third descriptive text and the first descriptive text in the target prompt word, the semantic coherence between the preceding and following video segments corresponding to the initial cumulative segmentation point is determined; When the semantic coherence degree represents the semantic coherence between the preceding and following video segments, the initial identifier of the initial cumulative segmentation point is deleted. When the semantic coherence degree indicates that the semantics between the preceding and following video segments are not coherent, the initial identifier of the initial cumulative segmentation point is determined as the second identifier.

6. The method according to claim 5, characterized in that, The calibration of the initial fragment description text based on the second identifier to obtain the fourth description text includes: The accumulated keyframe sequence is segmented according to the initial accumulated segmentation point corresponding to the second identifier to obtain multiple second video segments; For each second video segment, when the second video segment includes multiple first video segments, the initial segment description texts of the multiple first video segments are semantically merged to obtain the fourth description text of the second video segment; When the second video segment includes one of the first video segments, the initial segment description text of the first video segment is determined as the fourth description text of the second video segment.

7. The method according to any one of claims 1 to 6, characterized in that, The step of updating the accumulated state information based on the third description text, the second identifier, and the fourth description text to obtain the accumulated state information for the next time window of the current time window includes: The third description text is appended to the first description text in the accumulated state information to obtain the merged description text; The merged description text is determined as the first description text in the cumulative status information of the next time window, the second identifier is determined as the first identifier in the cumulative status information of the next time window, and the fourth description text is determined as the second description text in the cumulative status information of the next time window.

8. The method according to any one of claims 1 to 6, characterized in that, The target video is segmented based on the accumulated state information obtained after iterative processing to obtain the video segmentation result of the target video, including: From the accumulated state information obtained after iterative processing, obtain the identifier sequence corresponding to the identifier of all segmentation points of the target video; Based on the identifiers in the identifier sequence, the target video is sequentially segmented to obtain multiple target video segments; The multiple target video segments are determined as the video segmentation results of the target video.

9. The method according to any one of claims 1 to 6, characterized in that, Before acquiring the current keyframe sequence and cumulative state information of the target video within the current time window, the method further includes: According to the preset sampling rate, the target video is frame-sampling to obtain an initial video frame sequence; Determine the scene embedding features and text information of each initial video frame in the initial video frame sequence; Based on the scene embedding features, the initial video frame sequence is segmented to obtain multiple video segments; Based on the scene embedding features and the text information, the initial video frames in each video segment are deduplicated to obtain the keyframe sequence of the target video.

10. The method according to claim 9, characterized in that, Based on the scene embedding features, the initial video frame sequence is segmented into multiple video segments, including: Determine the feature similarity between the scene embedding features of each pair of adjacent initial video frames in the initial video frame sequence; The feature similarity is normalized to obtain the normalized feature similarity. When the normalized feature similarity is less than the preset first feature similarity threshold, the later initial video frame among the two adjacent initial video frames is determined as the split-scene cutting point. Based on the segmentation points, the initial video frame sequence is segmented to obtain multiple video segments.

11. The method according to claim 10, characterized in that, The normalization process for the feature similarity to obtain the normalized feature similarity includes: Extract a sequence of reference video frames of a preset duration from the target video; Determine the reference feature similarity between the scene embedding features of each of the two adjacent reference video frames in the reference video frame sequence; Determine the mean similarity and standard deviation of the similarity of the multiple reference features; For each pair of adjacent initial video frames in the initial video frame sequence, the difference between the feature similarity and the mean similarity is determined. The ratio of the difference to the standard deviation of the similarity is determined as the normalized feature similarity.

12. The method according to claim 9, characterized in that, The process of deduplicating the initial video frames in each video scene based on the scene embedding features and the text information to obtain the keyframe sequence of the target video includes: For each video segment, when any two adjacent initial video frames in the video segment meet at least one of the following preset conditions, the text quality of the text information of each of the two adjacent initial video frames is determined; wherein, the preset conditions are: the feature similarity between the scene embedding features of each of the two adjacent initial video frames in the video segment is greater than a preset second feature similarity threshold; and the text similarity between the text information of each of the two adjacent initial video frames is greater than a preset text similarity threshold. When the text quality of the first initial video frame in any two adjacent initial video frames is greater than or equal to a preset quality threshold, the first initial video frame is retained. When the text quality corresponding to the preceding initial video frame is less than the quality threshold, the initial video frame corresponding to the highest text quality among any two adjacent initial video frames is retained. The retained initial video frames are sorted in chronological order to obtain the keyframe sequence of the target video.

13. A training method for a video understanding model, characterized in that, The method includes: According to the preset training time window, the following iterative processing is performed on the keyframes of the sample videos in the training sample set: The sample video is acquired within the current training time window, including the current keyframe sequence, cumulative state information, and target output truth value. The cumulative state information includes: a first description text for each historical keyframe in the historical keyframe sequence preceding the current training time window, a first identifier for the historical segmentation point corresponding to the historical keyframe sequence, and a second description text for each video segment obtained after segmentation according to the first identifier. The target output truth value includes: a frame description truth value sequence for the keyframes of the sample video, a truth value sequence for the identifier of the actual segmentation point of the sample video, and a segment description truth value sequence for the video segments obtained after segmenting the sample video according to the actual segmentation point. The first description text of the sample is subjected to noise injection processing to obtain noisy first description text; the first identifier of the sample is subjected to noise injection processing to obtain noisy first identifier; the second description text of the sample is subjected to noise injection processing to obtain noisy second description text. Based on the current keyframe sequence of the sample, the noisy first descriptive text, the noisy first identifier, and the noisy second descriptive text, a prediction result is determined using the video understanding model to be trained. The prediction result includes: a predicted third descriptive text for each current keyframe in the current keyframe sequence of the sample, a predicted second identifier for the cumulative segmentation point corresponding to the cumulative keyframe sequence of the sample, and a predicted fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence of the sample according to the cumulative segmentation point. The cumulative keyframe sequence of the sample includes the current keyframe sequence of the sample and the historical keyframe sequence of the sample. Based on the predicted third description text, the predicted second identifier, and the predicted fourth description text, the sample cumulative state information is updated to obtain the sample cumulative state information for the next training time window of the current training time window; the current training time window and the next training time window partially overlap in the time dimension. Based on the predicted third description text and the frame description truth sequence, the predicted second identifier and the identifier truth sequence, and the predicted fourth description text and the fragment description truth sequence, loss calculation is performed to obtain a multi-task joint loss value; Based on the multi-task joint loss value, the network parameters of the video understanding model to be trained are updated to obtain the trained video understanding model.

14. The method according to claim 13, characterized in that, The loss calculation based on the predicted third descriptive text and the frame description truth sequence, the predicted second identifier and the identifier truth sequence, and the predicted fourth descriptive text and the fragment description truth sequence, to obtain a multi-task joint loss value, includes: From the frame description truth value sequence, obtain the target frame description truth value for the current key frame of each sample; From the true value sequence of the identifiers, obtain the true value of the target identifier for the target true segmentation point corresponding to the sample accumulated keyframe sequence; From the fragment description truth value sequence, obtain the target fragment description truth value of each sample video fragment obtained by segmenting the sample cumulative keyframe sequence according to the target true segmentation point; The multi-task joint loss value is obtained by calculating the loss based on the predicted third descriptive text and the target frame description ground value, the predicted second identifier and the target identifier ground value, and the predicted fourth descriptive text and the target fragment description ground value.

15. The method according to claim 14, characterized in that, The multi-task joint loss value is obtained by performing loss calculation based on the predicted third descriptive text and the target frame description ground truth, the predicted second identifier and the target identifier ground truth, and the predicted fourth descriptive text and the target fragment description ground truth, including: Based on the predicted third descriptive text and the target frame description truth value, a first loss calculation is performed to obtain the key frame description loss value; A second loss is calculated based on the predicted second identifier and the true value of the target identifier to obtain the video segmentation loss value; The third loss is calculated based on the predicted fourth description text and the target fragment description truth value to obtain the fragment description loss value; Based on the prediction results and the target output true value, a fourth loss calculation is performed to obtain the global loss value; Based on preset first, second, third, and fourth weights, the keyframe description loss value, the video segmentation loss value, the segment description loss value, and the global loss value are weighted and summed to obtain the multi-task joint loss value.

16. A video segmentation device, characterized in that, The device includes: The first iterative processing module is configured to perform the following iterative processing on keyframes of a target video according to a preset time window: acquiring the current keyframe sequence and accumulated state information of the target video within the current time window; the accumulated state information includes: a first description text for each historical keyframe in a historical keyframe sequence preceding the current time window, a first identifier of a historical segmentation point corresponding to the historical keyframe sequence, and a second description text for each video segment obtained after segmenting the historical keyframe sequence according to the historical segmentation point; based on the current keyframe sequence, the first description text, the first identifier, and the second description text... The text defines a third descriptive text for each current keyframe in the current keyframe sequence, a second identifier for the cumulative segmentation point corresponding to the cumulative keyframe sequence, and a fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence according to the cumulative segmentation point; the cumulative keyframe sequence includes the current keyframe sequence and the historical keyframe sequence; based on the third descriptive text, the second identifier, and the fourth descriptive text, the cumulative state information is updated to obtain the cumulative state information of the next time window of the current time window; the current time window and the next time window partially overlap in the time dimension; The video segmentation module is used to segment the target video based on the accumulated state information obtained after the iterative processing of all keyframes in the target video, and obtain the video segmentation result of the target video.

17. A training device for a video understanding model, characterized in that, The device includes: The second iterative processing module is used to perform the following iterative processing on the sample keyframes of the sample videos in the training sample set according to a preset training time window: obtaining the current keyframe sequence of the sample videos within the current training time window, the sample cumulative state information, and the target output truth value; the sample cumulative state information includes: a sample first description text for each sample historical keyframe in the sample historical keyframe sequence located before the current training time window, a sample first identifier for the sample historical segmentation point corresponding to the sample historical keyframe sequence, and a sample second description text for each sample video segment obtained after segmentation according to the sample first identifier. The target output truth value includes: a frame description truth value sequence for the sample keyframes of the sample video, a truth value sequence for the identifiers of the actual segmentation points of the sample video, and a segment description truth value sequence for the sample video segments obtained after segmenting the sample video according to the actual segmentation points; noise injection processing is performed on the first sample description text to obtain a noisy first description text; noise injection processing is performed on the first sample identifier to obtain a noisy first identifier; noise injection processing is performed on the second sample description text to obtain a noisy second description text; based on the current keyframe sequence of the sample, the noisy first description text, and the noisy second description text... A first identifier and the noisy second descriptive text are used to determine a prediction result through a video understanding model to be trained. The prediction result includes: a predicted third descriptive text for each current keyframe in the current keyframe sequence of the sample, a predicted second identifier for the cumulative segmentation point of the cumulative keyframe sequence of the sample, and a predicted fourth descriptive text for each video segment obtained after segmenting the cumulative keyframe sequence of the sample according to the cumulative segmentation point. The cumulative keyframe sequence of the sample includes the current keyframe sequence of the sample and the historical keyframe sequence of the sample. Based on the predicted third descriptive text, the predicted second identifier, and the predicted third descriptive text, a prediction result is determined. The fourth descriptive text is tested, and the cumulative sample state information is updated to obtain the cumulative sample state information for the next training time window of the current training time window; the current training time window and the next training time window partially overlap in the time dimension; loss is calculated based on the predicted third descriptive text and the frame description truth sequence, the predicted second identifier and the identifier truth sequence, and the predicted fourth descriptive text and the segment description truth sequence to obtain a multi-task joint loss value; based on the multi-task joint loss value, the network parameters of the video understanding model to be trained are updated to obtain the trained video understanding model.

18. An electronic device, characterized in that, include: Memory is used to store executable instructions or computer programs. The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the video segmentation method according to any one of claims 1 to 12, or implements the training method for the video understanding model according to any one of claims 13 to 15.

19. A computer-readable storage medium, characterized in that, The device stores computer-executable instructions or computer programs for inducing a processor to execute the computer-executable instructions or computer programs to implement the video segmentation method according to any one of claims 1 to 12, or to implement the training method for the video understanding model according to any one of claims 13 to 15.

20. A computer program product, characterized in that, The computer program product includes computer-executable instructions or a computer program, which, when executed by a processor, implement the video segmentation method according to any one of claims 1 to 12, or implement the training method for the video understanding model according to any one of claims 13 to 15.