Method and device for splitting video and electronic equipment

Through the dual analysis of speech recognition and picture content scoring, the problem of poor video segmentation effect is solved, and the accurate segmentation and structuring of video content is achieved.

CN120769112APending Publication Date: 2025-10-10MIAOZHEN INFORMATION TECHNOLOGY (ZIYANG) CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511083877.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies are unable to perform in-depth analysis of video content, resulting in poor video splitting effects.

Method used

The original video is split through speech recognition technology, the picture content score of each frame is calculated, the target frame is determined according to the speech and picture content scores, and the video segments are cut based on this to generate multiple video storyboards with independent visual content.

Benefits of technology

It achieves accurate segmentation of video content and improves the structuring of video content, making it into multiple segments and storyboards with clear semantics and visual independence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769112A_ABST
    Figure CN120769112A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video processing, and discloses a method and device for splitting a video and electronic equipment, and the method comprises the steps: carrying out the voice recognition of an original video, splitting the original video according to a voice recognition result, and obtaining video segments; calculating a picture content score of each frame in each video segment; determining a plurality of target frames according to the picture content score of each frame; and according to each target frame, segmenting the video segment to which the target frame belongs to obtain a plurality of video sub-mirrors. Through double analysis of voice recognition and picture content score, the video content is split into a plurality of fragments and sub-mirrors with definite semantic and visual independence, the structured degree of the video content is improved, and accurate splitting of the video content is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, for example, to a method and device for splitting videos, and an electronic device. Background Art

[0002] With the rapid development of video content, video has become a crucial medium for brand promotion, product marketing, and user interaction. Brands use videos to showcase product selling points and convey brand value, while users want to quickly grasp the core content of the video. Therefore, developing technology that can accurately identify and segment videos is crucial for improving user experience and commercial value.

[0003] In the related art, a video processing method is disclosed, which relies on a simple editing tool or extracts video clips based on user behavior data.

[0004] During the implementation of the embodiments of the present disclosure, it was found that at least the following problems exist in the related art:

[0005] The related technology cannot perform in-depth analysis of video content, and the effect of video splitting is poor.

[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0007] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0008] The embodiments of the present disclosure provide a method and apparatus for splitting a video, and an electronic device, so as to accurately split the video content.

[0009] In some embodiments, the method for splitting a video includes: performing speech recognition on the original video, and splitting the original video according to the speech recognition results to obtain video segments; calculating the picture content score of each frame in each video segment; determining multiple target frames according to the picture content score of each frame; and according to each target frame, dividing the video segment to which the target frame belongs to obtain multiple video storyboards.

[0010] Optionally, the original video is split according to the speech recognition results to obtain video segments, including: obtaining each text in the speech recognition results, and the start timestamp and end timestamp of each text in the original video timeline; merging multiple adjacent texts with contextual connections to obtain multiple merged texts; and splitting the original video according to the start timestamp of the first text and the end timestamp of the last text in each merged text to obtain multiple video segments.

[0011] Optionally, the method for splitting a video further includes: when there is no text in the speech recognition result, not splitting the original video, and treating the original video as a video segment.

[0012] Optionally, the picture content score of each frame in each video segment is calculated, including: calculating the average brightness, average saturation and edge strength of each frame; determining the brightness weight, saturation weight and edge weight of the current frame based on the contribution of the brightness, saturation and edge changes in the current frame to the video content change; and calculating the picture content score of the current frame based on the average brightness and brightness weight, average saturation and saturation weight, edge strength and edge weight of the current frame.

[0013] Optionally, multiple target frames are determined based on the picture content score of each frame, including: classifying all picture content scores to obtain maximum value data, intermediate value data and extremely low value data; determining a segmentation threshold based on a preset algorithm; the segmentation threshold is lower than the minimum value of the intermediate value data; when the target picture content score exceeds the segmentation threshold, determining the frame corresponding to the target picture content score as the target frame.

[0014] Optionally, the segmentation threshold is determined based on a preset algorithm, including: calculating the first average value of the intermediate value data and the second average value of the extremely low value data, and calculating the difference between the first average value and the second average value; determining the difference of a preset ratio as a buffer value; calculating the difference between the lowest value of the intermediate value data and the buffer value as the segmentation threshold.

[0015] Optionally, the method for splitting a video also includes: when the duration of the current video frame is less than the first duration, merging the current video frame with the next video frame; or, when the duration of the current video frame is greater than the second duration, continuing to split the current video frame; or, when the correlation between the last video frame of the current video segment and the first video frame of the next video segment exceeds a set threshold, merging the two into one video frame.

[0016] Optionally, the method for splitting a video further comprises: after obtaining the plurality of video shots, extracting multi-modal feature information in each video shot; and based on the multi-modal large language model, performing label content extraction according to the multi-modal feature information of each video shot, to realize labeling of each video shot.

[0017] In some embodiments, the apparatus for splitting a video comprises a processor and a memory storing program instructions, the processor being configured to execute a method for splitting a video as described when running the program instructions.

[0018] In some embodiments, the electronic device comprises: an electronic device body; and an apparatus for splitting a video as described, installed on the electronic device body.

[0019] The method and apparatus for splitting a video, and the electronic device provided by the embodiments of the present disclosure can achieve the following technical effects:

[0020] In the embodiments of the present disclosure, the original video is split by voice recognition technology, and the video can be divided into multiple segments with clear themes according to the semantic changes of voice content. Then, the picture content score of each frame in each video segment is calculated, which can accurately identify the degree of change of picture content in the video. According to the picture content score, multiple target frames are determined, and the video segments are cut based on the target frames, which can generate multiple video shots with independent visual content. Therefore, by double analysis of voice recognition and picture content score, the video content is split into multiple segments and shots with clear semantics and visual independence, which improves the structural degree of video content and realizes accurate splitting of video content.

[0021] The foregoing general description and the following description are merely exemplary and explanatory, and are not intended to limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0022] One or more embodiments are exemplarily illustrated by corresponding drawings, which are not intended to limit the embodiments, elements with the same reference numerals in the drawings are shown as similar elements, the drawings do not constitute a proportional limit, and wherein:

[0023] Figure 1 is a schematic diagram of a method for splitting a video provided by the embodiments of the present disclosure;

[0024] Figure 2 is a schematic diagram of another method for splitting a video provided by the embodiments of the present disclosure;

[0025] Figure 3 is a schematic diagram of another method for splitting a video provided by the embodiments of the present disclosure;

[0026] Figure 4 is a schematic diagram of picture content scores of each frame in a video provided by an embodiment of the present disclosure;

[0027] Figure 5 is a schematic diagram of another method for splitting a video provided by an embodiment of the present disclosure;

[0028] Figure 6 is a schematic diagram of an apparatus for splitting a video provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] In order to enable a person skilled in the art to more fully understand the features and technical contents of the embodiments of the present disclosure, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, which are used only for reference and are not intended to limit the embodiments of the present disclosure. In the following technical description, in order to facilitate explanation, a plurality of details are provided to provide a full understanding of the disclosed embodiments. However, one or more embodiments can still be implemented without these details. In other cases, well-known structures and devices can be simplified to facilitate the drawings.

[0030] The terms "first", "second", and the like in the technical solutions described in the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion.

[0031] Unless otherwise specified, the term "a plurality of" means two or more.

[0032] In the embodiments of the present disclosure, the character " / " represents an "or" relationship between the objects before and after it. For example, A / B represents: A or B.

[0033] The term "and / or" is a description of the association between objects, which means that there can be three relationships. For example, A and / or B means: A or B, or, A and B, the three relationships.

[0034] The term "corresponding" can refer to an association or binding relationship. A and B correspond to each other means that there is an association or binding relationship between A and B.

[0035] In combination Figure 1 As shown, the embodiments of the present disclosure provide a method for splitting a video, the execution subject of the method can be a processor, and the method comprises:

[0036] S101, the processor performs speech recognition on the original video, and splits the original video according to the speech recognition result to obtain a video segment.

[0037] S102: The processor calculates the picture content score of each frame in each video segment.

[0038] S103: The processor determines multiple target frames according to the picture content score of each frame.

[0039] S104: The processor divides the video segment to which each target frame belongs according to the target frame to obtain multiple video segments.

[0040] In the embodiment of the present disclosure, the original video is split by speech recognition technology, and the video can be divided into multiple segments with clear themes according to the semantic changes of the speech content. Then, the picture content score of each frame in each video segment is calculated, and the degree of change of the picture content in the video can be accurately identified. According to the picture content score, multiple target frames are determined, and the video segment is divided based on this, so that multiple video storyboards with independent visual content can be generated. Therefore, through the dual analysis of speech recognition and picture content score, the video content is split into multiple segments and storyboards with clear semantics and visual independence, which improves the structuring of the video content and achieves accurate segmentation of the video content.

[0041] Optionally, the original video is split according to the speech recognition results to obtain video segments, including: obtaining each text in the speech recognition results, and the start timestamp and end timestamp of each text in the original video timeline; merging multiple adjacent texts with contextual connections to obtain multiple merged texts; and splitting the original video according to the start timestamp of the first text and the end timestamp of the last text in each merged text to obtain multiple video segments.

[0042] Combine Figure 2 As shown, the embodiment of the present disclosure provides another method for splitting a video, including:

[0043] S201: The processor performs speech recognition on the original video, and obtains each text in the speech recognition result, as well as the start timestamp and end timestamp of each text in the timeline of the original video.

[0044] S202: The processor merges multiple adjacent texts that have contextual connections to obtain multiple merged texts.

[0045] S203: The processor splits the original video according to the start timestamp of the first text and the end timestamp of the last text in each merged text to obtain multiple video segments.

[0046] S204: The processor calculates the picture content score of each frame in each video segment.

[0047] S205: The processor determines multiple target frames according to the picture content score of each frame.

[0048] S206: The processor divides the video segment to which each target frame belongs according to the target frame to obtain multiple video segments.

[0049] In this embodiment, by obtaining the text and its corresponding timestamp in the speech recognition result and merging the text with contextual connection, the original video can be accurately split according to the timestamp range of the merged text. The semantic integrity and coherence of each video segment are ensured, and the semantic break caused by simple time segmentation is avoided. For example, for [text 1, text 1 start timestamp, text 1 end timestamp], [text 2, text 2 start timestamp, text 2 end timestamp], after merging text 1 and text 2, [text 1 + text 2, text 1 start timestamp, text 2 end timestamp] is obtained.

[0050] Optionally, the method for splitting a video further includes: when there is no text in the speech recognition result, not splitting the original video, and treating the original video as a video segment.

[0051] In this embodiment, for videos without voice content, such as pure music, the original video is directly segmented as a whole, avoiding the loss or incorrect segmentation of video content due to the lack of text information. This ensures the integrity of the video content, and even in the absence of voice information, the original video can still be fully preserved and processed.

[0052] Optionally, the picture content score of each frame in each video segment is calculated, including: calculating the average brightness, average saturation and edge strength of each frame; determining the brightness weight, saturation weight and edge weight of the current frame based on the contribution of the brightness, saturation and edge changes in the current frame to the video content change; and calculating the picture content score of the current frame based on the average brightness and brightness weight, average saturation and saturation weight, edge strength and edge weight of the current frame.

[0053] In this embodiment, the calculation of the picture content score can be implemented by the Python video picture detection library PySceneDetect. By calculating the average brightness, average saturation and edge strength of each frame, the visual features of each frame can be accurately analyzed to capture the key visual changes in the video. According to the contribution of the brightness, saturation and edge changes in the current frame to the change in video content, it can adapt to different video content and ensure the rationality of weight distribution, thereby improving the accuracy of the picture content score. For example, if the brightness change between adjacent frames is large, then the brightness weight should be higher; if the saturation change between adjacent frames is large, then the saturation weight should be higher. The degree of change in the picture content score of adjacent frames can reflect the degree of change in the content of the picture.

[0054] Optionally, calculating the average brightness of each frame includes: converting each frame into a grayscale image, and then calculating the average of all pixel values ​​to determine the average brightness.

[0055] Optionally, calculating the average saturation of each frame includes: converting each frame into an HSV color space, and then calculating an average value of a saturation channel to determine the average saturation.

[0056] Optionally, calculating the edge strength of each frame includes: using an edge detection algorithm to calculate the edge strength of each frame.

[0057] Optionally, multiple target frames are determined based on the picture content score of each frame, including: classifying all picture content scores to obtain maximum value data, intermediate value data and extremely low value data; determining a segmentation threshold based on a preset algorithm; the segmentation threshold is lower than the minimum value of the intermediate value data; when the target picture content score exceeds the segmentation threshold, determining the frame corresponding to the target picture content score as the target frame.

[0058] Combine Figure 3 As shown, the embodiment of the present disclosure provides another method for splitting a video, including:

[0059] S301: The processor performs speech recognition on the original video and splits the original video according to the speech recognition result to obtain video segments.

[0060] S302: The processor calculates the picture content score of each frame in each video segment.

[0061] S303: The processor classifies all the picture content scores to obtain maximum value data, intermediate value data, and extremely low value data.

[0062] S304: The processor determines a segmentation threshold based on a preset algorithm; the segmentation threshold is lower than the minimum value of the intermediate value data.

[0063] S305 : When the target picture content score exceeds the segmentation threshold, the processor determines that the frame corresponding to the target picture content score is the target frame.

[0064] S306: The processor divides the video segment to which each target frame belongs according to the target frame to obtain multiple video segments.

[0065] In this embodiment, by classifying the picture content scores and determining the segmentation threshold, the frames in the video with significant picture content changes can be accurately identified as target frames. Figure 4 As shown in FIG, in a video, for the picture content scores of different frames, the picture content scores in [0, 32] are classified as extremely low value data, the picture content scores in (32, 70) are classified as intermediate value data, and the picture content scores in [70, 100] are classified as maximum value data.

[0066] Optionally, all the screen content scores are classified, including: classifying all the screen content scores based on a data clustering method, such as k-means, DBSCAN and other methods.

[0067] Optionally, the segmentation threshold is determined based on a preset algorithm, including: calculating the first average value of the intermediate value data and the second average value of the extremely low value data, and calculating the difference between the first average value and the second average value; determining the difference of a preset ratio as a buffer value; calculating the difference between the lowest value of the intermediate value data and the buffer value as the segmentation threshold.

[0068] In this embodiment, by calculating the average of the median data and the extreme low data, and considering the difference of a preset ratio as a buffer value, the segmentation threshold can be accurately determined, which helps to more accurately identify key frames in the video. The segmentation threshold can be adjusted according to the score distribution of different video content, making the determination of the segmentation threshold more robust and capable of processing various types of videos. For example, if the average of the median data is 50 and the average of the extreme low data is 10, then the difference between the two is 40. In this case, 40% of the difference is selected as the buffer value, that is, the buffer value is 16, and the segmentation threshold is determined to be 34.

[0069] Optionally, the preset ratio is in the range of [30%, 50%]. More specifically, the preset ratio is 30%, 35%, 40%, 45% or 50%.

[0070] Optionally, according to each target frame, the video segment to which the target frame belongs is divided to obtain multiple video storyboards, including: taking the starting frame of the video segment as the starting frame of the first video storyboard, and taking the previous frame of the first target frame as the ending frame of the first video storyboard; taking the current target frame as the starting frame of the current video storyboard, and taking the previous frame of the next target frame as the ending frame of the current video storyboard; taking the last target frame as the starting frame of the last video storyboard, and taking the ending frame of the video segment as the ending frame of the last video storyboard.

[0071] Optionally, the method for splitting a video also includes: when the duration of the current video frame is less than the first duration, merging the current video frame with the next video frame; or, when the duration of the current video frame is greater than the second duration, continuing to split the current video frame; or, when the correlation between the last video frame of the current video segment and the first video frame of the next video segment exceeds a set threshold, merging the two into one video frame.

[0072] In this embodiment, by merging short shots and splitting long shots, the duration of the video shots can be optimized to better fit within a preset duration range, more clearly displaying changes in the video content, and helping to improve the coherence and comprehensibility of the video content. In addition, by analyzing adjacent shots of adjacent video segments and merging highly correlated shots, the coherence of the video segments can be optimized, making the video shots more consistent with the actual logic of the video content.

[0073] Optionally, the first duration is in the range of [1s, 3s]. More specifically, the first duration is 1s, 2s, or 3s.

[0074] Optionally, the second duration is in the range of [15s, 20s]. More specifically, the first duration is 15s, 16s, 17s, 18s, 19s, or 20s.

[0075] Optionally, the method for splitting a video also includes: after obtaining multiple video storyboards, extracting multimodal feature information in each video storyboard; based on a multimodal large language model, extracting label content according to the multimodal feature information of each video storyboard to achieve labeling of each video storyboard.

[0076] Combine Figure 5 As shown, the embodiment of the present disclosure provides another method for splitting a video, including:

[0077] S501: The processor performs speech recognition on the original video and splits the original video according to the speech recognition result to obtain video segments.

[0078] S502: The processor calculates the picture content score of each frame in each video segment.

[0079] S503: The processor determines multiple target frames according to the picture content score of each frame.

[0080] S504: The processor divides the video segment to which each target frame belongs according to the target frame to obtain multiple video segments.

[0081] S505: The processor extracts multimodal feature information from each video frame.

[0082] S506: The processor extracts label content based on the multimodal large language model and the multimodal feature information of each video storyboard to achieve labeling of each video storyboard.

[0083] In this embodiment, the video content is analyzed from multiple dimensions within each video frame, extracting multimodal feature information that comprehensively reflects the characteristics of the video content. Finally, label content is extracted from this multimodal feature information based on a multimodal large language model. This allows for a more accurate understanding of the video content, flexibly adapting to different video content and annotation requirements, and achieving higher-precision video frame labeling.

[0084] Optionally, in each video storyboard, multimodal feature information is extracted, including: obtaining target text information and target visual information of the target video storyboard; integrating the target text information and the target visual information to obtain multimodal feature information of the target video storyboard.

[0085] In this embodiment, multimodal feature information such as text information and visual information is extracted in each video storyboard, which can more comprehensively reflect the characteristics of the video content. Integrating the target text information and the target visual information can establish a semantic association between the text and visual information, thereby adapting to different types of video content. Whether it is a video mainly based on dialogue or a video mainly based on visual effects, it can provide an accurate annotation basis by extracting and integrating multimodal information. The integration of multimodal feature information can reduce the ambiguity and errors that may be caused by single modal information. For example, relying solely on visual information may not accurately judge a character's emotions, but combining text information (such as lines) can provide a more accurate judgment.

[0086] Optionally, obtaining target text information of the target video storyboard includes: performing speech recognition on the target video storyboard to convert the audio content into initial text; performing semantic analysis on the initial text to extract key information as target text information.

[0087] In this embodiment, speech recognition and semantic analysis enable more accurate extraction of textual information from videos. Semantic analysis provides a deeper understanding of the text and extracts more valuable information. For example, it not only extracts the actual content of the dialogue, but also analyzes its emotional tendencies and contextual relationships, improving the accuracy and depth of subsequent video annotation. Furthermore, when acquiring target textual information, it can also extract subtitles from the target video storyboard to fully capture the language content of the video.

[0088] Optionally, obtaining target visual information of a target video storyboard includes: taking screenshots at the start position, middle position, and end position of the video storyboard according to the duration of the video storyboard to obtain screenshot information; and performing image analysis on the screenshot information to obtain target visual information.

[0089] In this embodiment, the screenshots at the start position, middle length position and end position of the video storyboard represent the main content of the entire storyboard, providing representative visual information of the video storyboard. Image analysis is performed on the screenshot information to extract target visual information. Image analysis may include object recognition, scene recognition, color analysis, texture analysis, etc. Object recognition can identify the main objects in the screenshot, such as characters, props, scenes, etc., providing key information for understanding the content of the video. Scene recognition can identify the scene type in the screenshot, such as indoor, outdoor, natural landscape, etc., to help understand the background and environment of the video. Color and texture analysis can analyze the color and texture features in the screenshot to provide a description of the visual style and atmosphere of the video.

[0090] Optionally, the multimodal large language model includes gemini-2.0-flash, qwen2.5-vl-72b-instr uct, and other multimodal large language models that can understand text modal content and image modal content.

[0091] Optionally, label content extraction is performed separately according to the multimodal feature information of each video storyboard, including: obtaining preset labels to be extracted; each label to be extracted includes a label name and a label description corresponding to the label name; among multiple label names corresponding to the labels to be extracted, determining the target label name of the current video storyboard according to the multimodal feature information; according to the label description corresponding to the target label name, performing content extraction in the multimodal feature information to generate label content of the current video storyboard under each target label name.

[0092] In this embodiment, through the preset label system to be extracted, the target label name of each video storyboard can be accurately determined, and content extraction is performed in the multimodal feature information according to the label description corresponding to the target label name, so that the label content of each video storyboard under each target label name can be efficiently generated.

[0093] Optionally, the tag name and tag description of the tag to be extracted are as shown in Table 1.

[0094] Table 1

[0095]

[0096]

[0097] Optionally, based on the tag description corresponding to the target tag name, tag content extraction is performed in the multimodal feature information, including: searching for relevant target feature information in the multimodal feature information based on the tag description; and generating tag content matching the tag description based on the target feature information.

[0098] In this embodiment, the tag description is used as a guide to accurately search for relevant target feature information in the multimodal feature information, ensuring the pertinence and accuracy of the extraction process. For example, for the "actor action" tag, not only the action description in the text is searched, but also the action details in the visual information are combined to ensure that the extracted information is comprehensive and accurate. Based on the searched target feature information, tag content that matches the tag description is generated to ensure the accuracy and consistency of the tag content. For example, for the "shot description" tag, the system combines text and visual information to generate a concise and clear description, such as "the character points to a tall mountain in the distance."

[0099] In a specific embodiment, assume that the multimodal feature information in a video storyboard includes: textual information such as the line "Look, that mountain is so high!"; and visual information such as the presence of a tall mountain in the scene, with a character pointing in its direction. By analyzing this multimodal feature information, it can be determined that the target tag for the current video storyboard is "shot description," and the tag description for "shot description" is "a brief description of the shot's content, limited to approximately thirty characters." Therefore, based on this tag description, the tag content generated for the current video storyboard under "shot description" is "a character pointing at a tall mountain in the distance."

[0100] Optionally, the method for splitting a video further includes: generating custom tags according to different application scenarios and application requirements; and updating preset tags to be extracted according to the custom tags.

[0101] In this embodiment, users can input the names and descriptions of custom tags through the system interface or configuration file according to different application scenarios and application requirements, so that the system can adapt to various specific video content analysis needs. For example, users may need to add a "special effects" tag to mark special effects scenes in the video; in advertising video analysis, users can add a "product display" tag; in movie analysis, users can add a "plot twist" tag. Based on the custom tag name and description entered by the user, it is stored in the tag library. These custom tags have the same structure as the preset tags, including tag name and tag description. By dynamically updating the extracted tags, it is ensured that the latest tag library is used in the video storyboard labeling process, which can flexibly respond to changing needs.

[0102] Optionally, after completing the labeling of each video storyboard, the storyboard information of the original video is determined.

[0103] In this embodiment, the frame information of the original video includes the video frames of the original video and the annotation information of each video frame.

[0104] Optionally, after obtaining the storyboard information of the original video, the target product data is obtained; a reference script is generated based on the target product data and the storyboard information of the original video; and the reference script is adjusted according to the adjustment parameters input by the user to generate the target script.

[0105] This embodiment incorporates the target product's category, brand, and selling point information, along with the original video's storyboard information, to generate a reference script that is highly relevant to the target product and consistent with the style and rhythm of the original video. This not only improves the efficiency and accuracy of script generation but also reduces creation costs. Furthermore, users can adjust the generated reference script to meet their needs, further improving the accuracy of the target script.

[0106] Optionally, a reference script is generated based on the target product data and the storyboard information of the original video, including: generating reference storyboard lines and reference shots based on the storyboard information of the original video; generating imitation content corresponding to the reference storyboard lines based on the target product data; and generating a reference script using the imitation content, reference storyboard lines, and reference shots.

[0107] In this embodiment, the style and rhythm of the original video are first analyzed based on its storyboard information, including the storyboard images, dialogues, shot descriptions, actors and actions, sets and props, scene and camera positions, and structure, to generate reference storyboard dialogues and reference shots. Then, based on the target product data, the reference storyboard dialogues are imitated to generate corresponding imitation content. Finally, a reference script is generated using the imitation content, reference storyboard dialogues, and reference shots.

[0108] In an optional embodiment, the target product data entered by the user includes: brand information for a cosmetics brand, category information for toys, and purchase information for fun. The main content of the original video is "a little girl imitating an adult putting on makeup in front of a mirror." Based on this content, a reference script can be generated. The generated reference script includes multiple storyboards, along with various script information such as the storyboard image, shot description, dialogue, brand placement, shot type, camera movement, duration, and actors appearing in each storyboard. Other script information may also be included, such as actor movements, expressions, dialogue emotions, set design requirements, background music, and costume design.

[0109] Optionally, the reference script is adjusted according to the adjustment parameters input by the user, including one or more of the following: in response to the user's instruction to delete the storyboard, a single storyboard in the reference script is deleted; in response to the user's instruction to modify the picture, the picture in the reference script is modified; in response to the user's intelligent optimization instruction, the content in the reference script is regenerated according to the context information in the reference script.

[0110] In this embodiment, users can adjust the reference script according to their needs, including deleting storyboards, modifying images, and optimizing content. This improves the user experience and enables the generation of a target script that better meets user expectations. Furthermore, users can instantly see the results of these adjustments, enabling them to quickly iterate and optimize script content. By using intelligent optimization capabilities to regenerate script content based on contextual information, the generated script content can be ensured to be not only relevant to the target product but also high-quality and coherent.

[0111] Optionally, the original video is obtained as follows: crawling the original video file according to the video link of the original video; or obtaining a local video file selected by the user and using the local video file selected by the user as the original video.

[0112] In this embodiment, users can choose to crawl videos from network links or directly upload local videos. This flexibility significantly improves the user experience and meets the needs of different users.

[0113] Optionally, according to the video link of the original video, the original video file is crawled, including: according to the video link of the original video, obtaining the title and thumbnail of the original video, and verifying according to the title and thumbnail; after the verification is passed, crawling the title information of the original video; after the title information is crawled, in response to the user's upload instruction, crawling the original video file and the interactive data of the original video.

[0114] In this embodiment, in the case of obtaining the original video through the video link, after the user fills in the video link on the User side and clicks confirm, the Server side obtains the title and thumbnail icon of the original video through the video link and verifies the video link. After verification, the title information is crawled using the Python service, and after successful crawling, the original video can be uploaded. If the link verification fails or the title crawling is unsuccessful, error information needs to be displayed for the user to adjust. In the case of multiple video links, after all the video links are verified, the user clicks batch upload on the User side, the Server side records the upload task information, calls the third-party interface to obtain the task_id, then crawls the original video file and the interactive data of the original video through the third-party supplier, and stores the video file and the interactive data in json format. In addition, the user can query the video upload progress through the third-party task_id, and after the video upload is completed, the Server side parses the json file of the video and the interactive information and updates the database. In the case of obtaining the original video through the local video file, after the user selects local upload on the User side and clicks upload, the Server side obtains the oss signature, and the User side uploads the local video file to the oss to store the local video file. After the user completes the upload of all local video files on the User side, the Server side stores the information of the local video file in the database.

[0115] After the original video is obtained, the script generation task is performed. After the user fills in the target commodity data and other script generation task information on the User side, the script generation is started. The Server side stores the task information, calls the agent robot, and obtains the script information. The agent robot can call the Dify analysis data to generate the script. After the Server side obtains and stores the script information, the algorithm is called to generate the split picture. In the split picture generation task queue, the Server side can call the algorithm to maintain the task state and automatically retry the task. After the picture generation task is completed, the Server side obtains the generated picture through the algorithm interface and stores the picture to the oss, and finally updates the script information and the task state.

[0116] In the method for generating a video script provided in an embodiment of the present disclosure, the concurrent number of algorithm interfaces of the task queue run by the processor does not exceed 3, and the call time of some interfaces is relatively long. For example, in a picture generation task, it takes 20s to 30s to generate a picture. In addition, a task may call multiple different algorithm interfaces, or call the same algorithm interface multiple times. For example, in a script generation task, multiple pictures need to be generated, and the total number of calls to the algorithm interface will be relatively large. The task queue of the embodiment of the present disclosure is formed using a producer-consumer model to control the maximum number of requests for the system to call a third-party interface, avoiding task failure caused by network request blocking. Use multi-threading to execute calling tasks, control the concurrent number of calls to third-party interfaces, and maximize performance. Add a retry mechanism to the queue, and retry a fixed number of times if the call fails. It will automatically fail after the number of retries is exhausted, avoiding failures caused by network fluctuations or system restarts.

[0117] Optionally, the method for splitting a video also includes: after obtaining a target script, generating multiple first storyboard videos according to the target script, and displaying the first storyboard videos on a video production page; in response to a user's editing operation on the first storyboard video on the video production page, obtaining a second storyboard video; in response to a user's synthesis operation on multiple second storyboard videos on the video production page, obtaining a target imitation video.

[0118] In this embodiment, after obtaining a target script, multiple first storyboard videos can be generated based on the target script and displayed on a video production page. After determining the first storyboard video, the first storyboard video can be processed to obtain a second storyboard video in response to a user's editing operation. In response to a user's synthesis operation on the multiple second storyboard videos, the multiple second storyboard videos can be synthesized to obtain a target imitation video, thereby reducing the difficulty of video production and improving the efficiency of video production.

[0119] Optionally, generating a plurality of first storyboard videos according to the target script includes: searching a pre-built material library according to the target script to determine a first storyboard video corresponding to the target script.

[0120] In this embodiment, a large amount of short video clips (i.e., storyboards) are included in the pre-built material library, each short video clip is provided with a label and ID according to classification (e.g., scene, emotion, and lens language), and the description text of each short video clip can be processed by a word segmentation tool, then converted into a text vector and associated with the short video clip for storage. When searching, the user can first determine the type, text vector, and label requirements of the imitation storyboard video according to the target script. Then, according to the type, text vector, and label requirements, fuzzy search is carried out in the material library by vector library (pgvector), determine the storyboard video that is the highest with the target script matching degree, and display it in the video production page as the first storyboard video. Or, determine a plurality of storyboard videos (e.g., 5 storyboard videos) that are ranked in front with the target script matching degree and display it in the video production page, and the first storyboard video is selected by the user in a plurality of storyboard videos.

[0121] Optionally, users can perform editing operations such as timeline editing (for example, cutting out segments, retaining 2 to 5 seconds of video segments from a 10-second storyboard video), effect enhancement (for example, adding filters, adjusting brightness or contrast, overlaying subtitles), and dynamic adjustment (for example, changing the playback speed to 0.5 times slow motion or 2 times acceleration, adding transition effects) on the first storyboard video in the video production page, and save the modified parts to obtain the second storyboard video.

[0122] Optionally, when responding to an editing operation, the first storyboard video after each editing operation can be saved to support undoing the editing operation of a certain step, avoiding the situation where the editing operation needs to be started again from the beginning after an error in the editing operation of a certain step.

[0123] Optionally, when performing the synthesis operation, the multiple second storyboard videos are multiple storyboards selected from the edited second storyboard video list, or all the storyboards in the list.

[0124] Optionally, when performing the synthesis operation, the target script can be subjected to sentiment analysis to determine the order of each second storyboard video, the transition effect (for example, flash cut or slide, etc.) and set other synthesis parameters (for example, background music, resolution and output format, etc.), so as to synthesize multiple second storyboard videos according to the determined parameters to generate the target imitation video.

[0125] Optionally, after the user enters a synthesis operation for multiple second-storyboard videos on the video production page, the synthesis operation request is sent to the backend server via the "POST / api / video / synthesis / create" request interface. The backend server receives the synthesis operation request through the "VideoSynthesisController" controller component and calls the "VideoSynthesisService" service layer component to process the request. The service layer component saves the basic information of the requested task in the "CmsTaskInfo" table, saves the task details in the "CmsTaskDetail" table, and asynchronously calls the "synthesizeVideoAsync" method to perform the following steps:

[0126] First, update the task status to "Downloading";

[0127] Second, use the "FileDownloadUtil" tool to download the video file to a local temporary directory;

[0128] Third, update the task status to "Synthesizing";

[0129] Fourth, call the "createFileList" method to create a synthetic file list;

[0130] Fifth, call the "concatVideos" method to use the FFmpeg tool to synthesize the video file into the target imitation video;

[0131] Sixth, update the task status to "Uploading";

[0132] Seventh, call the "uploadSynthesisVideo" method to upload the target simulation video to the OSS (Object Storage Service) platform;

[0133] Eighth, update the task status to "Completed" and store the OSS ID in the resultFileIds field;

[0134] Ninth, clean up temporary files during the synthesis process.

[0135] Optionally, if any step in the synthesis process fails, the task status is updated to "FAILED (error or exception)" and detailed error information is recorded in the errorMessage field to facilitate subsequent troubleshooting.

[0136] Optionally, during the execution of the synthesis operation, the user can query the task state and result of the synthesis operation task from the backend server through a "GET / api / video / synthesis / detail / {taskId}" request interface.

[0137] Optionally, in response to the user's editing operation on the first split video in the video production page, the second split video is obtained, including: determining the user's editing requirement according to the user's editing operation on the first split video; and adjusting one or more of the subtitle setting, the voiceover setting, the video audio mixing and the transition effect in the first split video according to the user's editing requirement to obtain the second split video.

[0138] Specifically, the editing operation on the video includes one or more of the subtitle setting, the voiceover setting, the video audio mixing and the transition effect.

[0139] Specifically, according to the editing operation on the first split video input by the user in the video production page, it can be determined how the user wants to edit the first split video. Therefore, the user's editing requirement can be determined.

[0140] Specifically, according to the user's editing requirement, it can be determined which one or which ones of the subtitle setting, the voiceover setting, the video audio mixing and the transition effect need to be performed. Therefore, the first split video can be adjusted according to the user's editing requirement to obtain the second split video.

[0141] Optionally, after each editing operation is performed, the effect after editing needs to be displayed on the video production page for the user to preview.

[0142] Optionally, when performing the editing operation, conflict detection is performed on each editing operation, and if there is a conflict, a prompt is given. For example, when the editing operation includes "mute original sound" and "enhance human voice", a conflict prompt can be given.

[0143] In this embodiment, after the user inputs the editing operation, the user's editing requirement can be automatically recognized, and the first split video is operated according to the editing requirement, which reduces the difficulty of the user editing the split video.

[0144] Optionally, if the first split video corresponding to the target script cannot be searched in the material library, it indicates that the material library does not store the first split video matching the target script. Therefore, in this case, the pre-constructed AI large model is called to generate the first split video matching the target script.

[0145] Optionally, based on a pre-built AI big model, a first storyboard video corresponding to the target script is generated, including: analyzing the target script to determine the script information, and creating a subject library based on the script information; wherein the subject library includes one or more pictures corresponding to the script information; filling the script information into the subject library, and adjusting the prompt words of the subject library according to the script information; and using the AI ​​big model to generate multiple first storyboard videos based on the pictures and script information in the subject library after the prompt words are adjusted.

[0146] Specifically, the target script is a text containing information such as scene descriptions, character dialogues, and action instructions. Therefore, by parsing the target script, key information (e.g., scene information: time, place, and characters; character information: character identity, appearance, and actions; plot information: plot, etc.) can be extracted to determine the script information.

[0147] Specifically, the types of pictures required for generating a storyboard video can be determined based on the script information. Therefore, a subject library including one or more pictures corresponding to the script information can be created based on the script information.

[0148] Specifically, after filling the script information into the subject library, the prompt words of the subject library are adjusted according to the script information. The script information can be embedded into the original prompt words of the subject library, so that the subsequent storyboard video generated based on the subject library is more in line with the intention of the script.

[0149] For example, the original prompt word in the subject library is "generate a running person", and according to the script information, the prompt word can be adjusted to "generate a protagonist wearing a black cloak running hard in the forest under the moonlight, with a movie-like picture style".

[0150] Specifically, for each shot, a matching picture is selected from the subject library, combined with the adjusted prompt words and script information and input into the AI ​​big model to generate the first storyboard video.

[0151] Optionally, users can also manually create a subject library in the video production page. The specific steps include: clicking the Create Subject Library option in the video production page; uploading one or more selected pictures, and filling in the name and label; analyzing the pictures to determine the style and description (which can be modified manually) as the prompt words for the subject library, and building the subject library.

[0152] Specifically, within a user's account, users can view and edit existing libraries on the video creation page. For example, they can select a library to view its details, modify the images, names, and tags within it, and re-analyze the style and description of images within the library. Alternatively, they can select one or more libraries to delete them, or select one or more libraries to generate a storyboard video.

[0153] In this embodiment, if the first storyboard video corresponding to the target script cannot be found in the pre-built resource library, the AI ​​large model will be used to generate the first storyboard video corresponding to the target script. This avoids the need for the user to manually collect the first storyboard video when the first storyboard video cannot be found in the resource library, further reducing the difficulty of video production and improving the efficiency of video production.

[0154] Optionally, the method for splitting a video further includes: after generating a first storyboard video corresponding to a target script, setting a label for the first storyboard video according to the target script; and saving the first storyboard video with the label set into a pre-built material library.

[0155] In this embodiment, if the first storyboard video is newly generated based on the AI ​​large model, the first storyboard video will be tagged according to the target, and then the tagged first storyboard video will be saved to the pre-built resource library. In this way, if the same or similar script is received again in the future, the relevant storyboard video can be directly retrieved from the resource library without having to generate it again, which improves the efficiency of video production.

[0156] Optionally, the method for splitting a video also includes: after determining the first storyboard video corresponding to the target script, displaying the target script and the first storyboard video on the video production page; in response to the user's modification operation on the target script, modifying the target script and updating the first storyboard video.

[0157] In this embodiment, by displaying the target script and the first storyboard video simultaneously on the video creation page, after the user modifies the target script, the updated first storyboard video can be displayed simultaneously. This makes it easier for users to intuitively compare the script and the storyboard effect, quickly identify problems, improve the user experience, and reduce the difficulty of video creation.

[0158] Optionally, the method for splitting a video further comprises: after obtaining the target imitation video, generating a sharing link of the target imitation video. The sharing link allows a user to directly access, play or download the video through a browser or other supported platform.

[0159] Optionally, after synthesizing the target imitation video, the user can generate a task request for a sharing link by inputting it in the video production page. The task request will be sent to the back-end server through the "POST / api / task / share / create" request interface and processed by the "TaskShareReportController" component. The "TaskShareReportController" component will first verify the task. The verification content includes: checking whether the task exists, checking whether the task belongs to the current user, checking the task type, and checking whether the video synthesis task has been completed. After the verification is passed, the "TaskShareReportController" component will generate sharing information (random authorization code, authorization code expiration time, user information and task details to be shared), build sharing content (build content containing the task name, sharer's name, task type, creation time and expiration time, and convert the content into a JSON string), save the sharing record and return the sharing link.

[0160] For example, the format of the sharing link is "baseUrl+" / api / open / share / "+authorizeCo de".

[0161] In this embodiment, by generating a sharing link (eg, URL) of the target imitation video, the user can share the target imitation video without directly transmitting a large video file, thereby improving the convenience of video sharing for the user.

[0162] Optionally, the method for splitting a video further includes: after obtaining the target imitation video, identifying highlight segments of the target imitation video.

[0163] Optionally, highlight segments of the target imitation video are identified, including: performing voice recognition on the target imitation video to obtain text data; performing image capture on the target imitation video to obtain video frame images; and identifying highlight segments of the target imitation video based on the text data and the video frame images to obtain highlight segments in the target imitation video.

[0164] In this embodiment, speech recognition is performed on the target imitation video and text data is obtained to convert the original audio data of the target imitation video into text data, providing a reliable data basis for the recognition of highlight segments. Then, image capture is performed on the target imitation video to obtain a video frame image. Finally, based on the text data and the video frame image, highlight segments of the target imitation video are recognized to obtain highlight segments in the target imitation video. By combining speech recognition with the image information provided by the captured video frame image, highlight segments of the target imitation video are recognized, eliminating the dependence on user behavior data in the highlight segment recognition process, and improving the accuracy of video content and highlight segment recognition.

[0165] Optionally, highlight segments of the target imitation video are identified based on the text data and the video frame images to obtain the highlight segments in the target imitation video, including: determining video content features based on the video frame images; determining highlight segment prompt words based on the text data and the video content features; and identifying highlight segments of the target imitation video based on the highlight segment prompt words and the video frame images to obtain the highlight segments in the target imitation video.

[0166] Optionally, the video content feature represents a feature in a video frame image that characterizes the characteristics of the video content. As an example, the video content feature can be a video content type, for example, a video content type of sports video, a video content type of film and television video, a video content type of fitness video, etc. As another example, the video content feature can also be a specific video content, for example, a video content feature can be a table tennis attack image in a sports video, a video content feature can also be a starting image of a track and field competition in a sports video, or a video content feature can also be a teacher's courseware image in a teaching video. It is understandable that the embodiments of the present disclosure only provide examples of some video content features, and the video content feature can be determined based on the specific video content.

[0167] Optionally, the highlight clip prompt word represents the natural language text to be input, which is used to identify the video highlight clip in the target imitation video to be input. As an example, in a sports video, the highlight clip prompt word can be the corresponding voice of the key sports action. For example, in a football game video, the highlight clip prompt word is shot, steal, corner kick, free kick, etc. In a table tennis game video, the highlight clip prompt word is forehand serve attack, backhand serve attack, spin serve, twist and pull, match point, serve change, etc. In a film and television video, the highlight clip prompt word can be the specific dialogue content at the climax of the plot in the film and television video.

[0168] In this embodiment, when identifying highlight segments in a target imitation video, the accuracy of the prompt word has a significant impact on the accuracy of the recognition results. By providing sufficient and reliable information from the video frame image and text data, and combining the highlight segment prompt word and video frame image determined by the video frame image and text data for recognition, the accuracy of video content and highlight segment recognition is further guaranteed.

[0169] Optionally, the video content feature includes a video content type.

[0170] Optionally, highlight clip prompt words are determined based on text data and video content features, including: when the video content type indicates that the video frame image matches the frame image corresponding to the sports video, determining the highlight clip prompt words as sports action words; or, when the video content type indicates that the video frame image matches the frame image corresponding to the film and television video, extracting the dialogue content of the plot peak segment of the target imitation video from the text data, and determining the highlight clip prompt words as the dialogue content of the plot peak segment; or, when the video content type indicates that the video frame image matches the frame image corresponding to the public welfare lecture video, extracting the summary narration or commentary of the target imitation video from the text data, and determining the highlight clip prompt words as the summary narration or commentary; or, when the video content type indicates that the video frame image matches the frame image corresponding to the cooking video, extracting the narration or commentary representing the cooking steps of the target imitation video from the text data, and determining the highlight clip prompt words as the narration or commentary representing the cooking steps.

[0171] In this embodiment, the highlight segment prompt words are customized in combination with the video content type to provide prompt words that match the video content type, thereby ensuring the accuracy of video content recognition and eliminating the dependence of highlight segment recognition on user behavior data.

[0172] Optionally, based on the highlight segment prompt words and the video frame image, the highlight segment of the target imitation video is identified to obtain the highlight segment in the target imitation video, including: performing feature extraction on the highlight segment prompt words to obtain a first feature vector corresponding to the highlight segment prompt words; performing feature extraction on the video frame image to obtain a second feature vector corresponding to the video data; obtaining the similarity between the first feature vector and the second feature vector; when the similarity is greater than or equal to a similarity threshold, determining that the segment corresponding to the video frame image is the highlight segment in the target imitation video.

[0173] Optionally, determining that the segment corresponding to the video frame image is the highlight segment in the target imitation video includes: extracting timestamps of video frame images whose similarity is greater than or equal to a similarity threshold; and merging multiple video frame images with consecutive timestamps to generate the highlight segment in the target imitation video. The multiple video frame images with consecutive timestamps indicate that the difference between the timestamps of the two video frame images is less than or equal to a time threshold. For example, the time threshold is greater than zero and less than or equal to 5 seconds.

[0174] In this embodiment, the highlight segments are extracted by extracting feature vectors and comparing similarities, which is beneficial for searching for the highlight segments of the target imitation video and improving the accuracy and reliability of highlight segment recognition.

[0175] Optionally, voice recognition is performed on the target imitation video to obtain text data, including: performing voice recognition on the target imitation video through automatic voice recognition technology to obtain voice recognition results and timestamps corresponding to the voice recognition results; wherein the voice recognition results include part or all of the dialogue, narration, and commentary; and constructing and generating text data based on the voice recognition results and the timestamps corresponding to the voice recognition results.

[0176] Optionally, performing image interception on the target imitation video to obtain a video frame image includes: intercepting a key frame of the target imitation video to obtain an initial video frame image; and performing image compression processing on the initial video frame image to obtain a video frame image.

[0177] The method for splitting videos provided by the embodiments of the present disclosure can accurately split video content and provide detailed storyboard information through automatic speech recognition technology and video analysis technology. The original video is divided into multiple segments and storyboards according to the degree of picture changes. This structured splitting method not only makes the organization of video content clearer and facilitates subsequent video editing and recommendation, but also meets the needs of brand merchants and video creators for refined content management, and provides strong support for in-depth analysis and efficient use of video content. In addition, through precise storyboard splitting, video creators can create and optimize content more efficiently, and video operators can conduct in-depth analysis based on segment information and storyboard information, optimize video recommendation strategies, and improve user retention, thereby improving the overall efficiency of video content creation and operation.

[0178] Combine Figure 6 As shown, an embodiment of the present disclosure provides an apparatus 600 for splitting a video, comprising a processor 601 and a memory 602. Optionally, the apparatus may further comprise a communication interface 603 and a bus 604. The processor 601, the communication interface 603, and the memory 602 may communicate with each other via the bus 604. The communication interface 603 may be used for information transmission. The processor 601 may invoke logic instructions in the memory 602 to execute the method for splitting a video according to the above embodiment.

[0179] In addition, the logic instructions in the memory 602 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0180] Memory 602, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present disclosure. Processor 601 executes the program instructions / modules stored in memory 602 to perform functional applications and data processing, thereby implementing the method for video segmentation in the above-mentioned embodiments.

[0181] The memory 602 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal device. Furthermore, the memory 602 may include high-speed random access memory and non-volatile memory.

[0182] An embodiment of the present disclosure provides an electronic device, comprising: an electronic device body, and the above-mentioned device for splitting a video. The device for splitting a video is installed on the electronic device body. The installation relationship described here is not limited to placement inside the electronic device, but also includes installation connections with other components of the electronic device, including but not limited to physical connections, electrical connections, or signal transmission connections. It can be understood by those skilled in the art that the device for splitting a video can be adapted to a feasible electronic device body, thereby realizing other feasible embodiments.

[0183] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned method for splitting a video.

[0184] The technical solutions of the embodiments of the present disclosure may be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or other media that can store program code.

[0185] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments only represent possible variations. Unless explicitly required, separate components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the technical solutions described in this application. As used in the technical solutions described in this application, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to also include plural forms. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups of these. In the absence of further restrictions, an element defined by the sentence "comprising a..." does not exclude the presence of other identical elements in the process, method or device that includes the element. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments can be referenced to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can be found in the description of the method part.

[0186] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0187] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical functional division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of the present disclosure may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0188] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operations of the systems, methods and computer program products according to the embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A method for splitting a video, characterized in that: include: Perform speech recognition on the original video and split the original video according to the speech recognition results to obtain video segments; Calculate the picture content score of each frame in each video segment; Determine multiple target frames based on the picture content score of each frame; According to each target frame, the video segment to which the target frame belongs is divided to obtain multiple video storyboards.

2. The method according to claim 1, characterized in that Split the original video according to the speech recognition results to obtain video segments, including: Get each text in the speech recognition result, as well as the start and end timestamps of each text in the original video timeline; Merging multiple adjacent texts with contextual connections to obtain multiple merged texts; The original video is split according to the start timestamp of the first text and the end timestamp of the last text in each merged text to obtain multiple video segments.

3. The method according to claim 2, characterized in that Also includes: When there is no text in the speech recognition result, the original video is not split and is treated as one video segment.

4. The method according to claim 1, wherein Calculate the image content score of each frame in each video segment, including: Calculate the average brightness, average saturation and edge intensity of each frame; Determine the brightness weight, saturation weight, and edge weight of the current frame according to the contribution of brightness, saturation, and edge changes in the current frame to the change of video content; The picture content score of the current frame is calculated based on the average brightness and brightness weight, average saturation and saturation weight, edge strength and edge weight of the current frame.

5. The method according to claim 1, wherein Based on the content score of each frame, multiple target frames are determined, including: Classify all screen content scores to obtain maximum value data, intermediate value data, and extremely low value data; Determine the segmentation threshold based on a preset algorithm; the segmentation threshold is lower than the minimum value of the median data; When the target picture content score exceeds the segmentation threshold, the frame corresponding to the target picture content score is determined as the target frame.

6. The method according to claim 5, characterized in that Determine the segmentation threshold based on a preset algorithm, including: Calculate a first average value of the median value data and a second average value of the extremely low value data, and calculate the difference between the first average value and the second average value; Determine a difference of a preset ratio as a buffer value; Calculate the difference between the lowest value of the intermediate data and the buffer value as the segmentation threshold.

7. The method according to any one of claims 1 to 6, characterized in that Also includes: When the duration of the current video frame is less than the first duration, merging the current video frame with the next video frame; or When the duration of the current video frame is greater than the second duration, continue to segment the current video frame; or When the correlation between the last video frame of the current video segment and the first video frame of the next video segment exceeds a set threshold, the two are merged into one video frame.

8. The method according to any one of claims 1 to 6, characterized in that Also includes: After obtaining multiple video storyboards, extract multimodal feature information from each video storyboard; Based on the multimodal large language model, label content is extracted according to the multimodal feature information of each video storyboard to achieve labeling of each video storyboard.

9. A device for splitting a video, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the method for splitting a video according to any one of claims 1 to 8 when running the program instructions.

10. An electronic device, characterized in that: include: Electronic device body; The apparatus for splitting a video according to claim 9 is installed in the electronic device body.

Citation Information

Cited By

  • Advertisement video picture segmentation method and device, storage medium and electronic device

    CN121126078A