Video segmentation method and device, and electronic equipment
By preprocessing and semantic understanding of the initial video, generating and combining intermediate short videos, the problem of inaccurate video scene segmentation in the prior art is solved, and the efficiency and creativity of video creation are improved.
Patent Information
- Application Number
- CN202510382112.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-25
AI Technical Summary
The existing video scene segmentation method is difficult to adapt to complex and changeable video content, resulting in poor video creation results.
By preprocessing the initial video, the intermediate short video is generated, the semantic understanding is performed, and the target short video is combined according to the semantic similarity of the intermediate short video.
It realizes accurate and reasonable segmentation under different lens switching, scene transformation and character subject transformation, and improves the efficiency and creativity of video creation.
Smart Images

Figure CN120378705A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of video processing. Specifically, the present application relates to a video segmentation method, apparatus, and electronic device. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the breakthrough of AIGC technology in the field of multimodal processing, the video creative industry has entered a new stage. The threshold for video creation has been lowered, and the richness, possibilities, and creativity have been significantly improved. In order to create more creative videos, video content understanding has received wide attention, and video scene segmentation is the key among them.
[0003] Video scene segmentation aims to divide videos according to scenes or time events, facilitating understanding and editing from different dimensions. Through scene segmentation, video clips can be flexibly combined to create diverse effects. Currently, video scene segmentation algorithms are mainly divided into three categories: shot boundary-based, action-based, and event-based segmentation. Traditional video scene segmentation methods often rely on specific features and rules and are difficult to adapt to complex and variable video content. Therefore, a more intelligent and efficient video scene segmentation method is needed. Summary of the Invention
[0004] In view of the problems existing in the above-mentioned prior art, embodiments of the present application provide a video segmentation method, apparatus, and electronic device, which perform short video editing operations based on video semantic understanding, thereby improving the accuracy, processing efficiency, and creativity of video segmentation operations.
[0005] In a first aspect, embodiments of the present application provide a video segmentation method, including the following steps:
[0006] Preprocess the initial video to generate multiple intermediate short videos;
[0007] Perform semantic understanding processing on the multiple intermediate short videos; and
[0008] Combine the multiple intermediate short videos according to the semantic similarity of the intermediate short videos to obtain a target short video.
[0009] Further, the combining the multiple intermediate short videos according to the semantic similarity of the intermediate short videos to obtain a target short video includes:
[0010] Combine adjacent intermediate short videos with related semantics according to the semantic similarity of the intermediate short videos to obtain a target short video.
[0011] Further, the combining adjacent intermediate short videos with related semantics according to the semantic similarity of the intermediate short videos to obtain a target short video includes:
[0012] Calculate the embedding vector of the intermediate short video; and
[0013] Judge the semantic similarity of the intermediate short videos according to the embedding vector, and combine adjacent intermediate short videos with related semantics to obtain a target short video.
[0014] Further, the step of judging the semantic similarity of the intermediate short videos according to the embedding vector, and combining adjacent intermediate short videos with related semantics to obtain a target short video includes:
[0015] When the embedding vector is greater than a preset threshold, it is determined that adjacent intermediate short videos are semantically similar, and adjacent intermediate short videos with related semantics are combined to obtain a target short video.
[0016] Further, the step of performing semantic understanding processing on the multiple intermediate short videos includes:
[0017] Perform semantic understanding processing on the multiple intermediate short videos through a video understanding module.
[0018] Further, the step of preprocessing the initial video to generate multiple intermediate short videos includes:
[0019] Generate multiple intermediate short videos by detecting transitions in the initial video.
[0020] Further, the step of generating multiple intermediate short videos by detecting transitions in the initial video includes:
[0021] Detect the transition positions of the initial video on the time axis of the initial video; and
[0022] Divide the initial video into the multiple intermediate short videos according to the transition boundaries of the transition positions of the initial video.
[0023] In a second aspect, an embodiment of the present application further provides a video segmentation device, including:
[0024] A preprocessing module, configured to preprocess an initial video to generate multiple intermediate short videos;
[0025] A semantic understanding module, configured to perform semantic understanding processing on the multiple intermediate short videos; and
[0026] A short video generation module, configured to combine the multiple intermediate short videos according to the semantic similarity of the intermediate short videos to obtain a target short video.
[0027] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is configured to implement the video segmentation method according to the first aspect described above when executing the program.
[0028] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program is used to implement the video segmentation method according to the first aspect described above.
[0029] The embodiments of the present application bring the following beneficial effects:
[0030] In the video segmentation method provided by the embodiments of the present application, first, the initial video is preprocessed to generate a plurality of intermediate short videos. Then, semantic understanding processing is performed on the plurality of intermediate short videos. Finally, according to the semantic similarity of the intermediate short videos, the plurality of intermediate short videos are combined to obtain the target short video. The video segmentation method provided by the embodiments of the present application can achieve accurate and reasonable video scene segmentation in different situations such as different shot transitions, different scene changes, character subject changes, and severe occlusions, avoiding splitting the same action, the same character subject, and the same event into different video segments, thereby affecting the effect of video creation. Thus, video scene segmentation is achieved in video creation, video editing, and the task of retrieving video segments with similar themes, improving the processing efficiency and creativity of related tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0032] Figure 1 It is a schematic flow framework diagram of the video segmentation method provided by the embodiments of the present application;
[0033] Figure 2 It is a schematic simulation flow diagram of the video segmentation method provided by the embodiments of the present application;
[0034] Figure 3 It is a structural block diagram of the video segmentation device provided by the embodiments of the present application;
[0035] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application.
[0036] The realization of the objectives of the present application, functional features, and advantages will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0038] In the description and claims of the present application and the above-mentioned accompanying drawings, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0039] Refer to Figure 1 and Figure 2 , Figure 1 and Figure 2 are the flow framework diagram and simulation flow schematic diagram of the video segmentation method according to the embodiments of the present application. As Figure 1 and Figure 2 shown, the video segmentation method according to the embodiments of the present application includes the following steps:
[0040] S101: Preprocess the initial video to generate a plurality of intermediate short videos;
[0041] Preprocessing is to pre-segment the initial video. Although the intermediate short videos are not the final segmentation results, they are consistent within each segment, with the time, location, subject, etc. remaining the same. Such a strategy for video preprocessing can reduce the computing power cost and time cost for the subsequent semantic understanding module, because in the subsequent semantic understanding module, it is not necessary to process each frame of the original video, but only need to process the pre-segmented short video segments.
[0042] S102: Perform semantic understanding processing on the plurality of intermediate short videos; and
[0043] Video semantic understanding aims to extract and understand language information from the video for in-depth understanding and analysis of the video. It not only involves the processing of multimedia information such as images, sounds, and texts in the video, but also includes semantic analysis and understanding of these information.
[0044] S103: Combine the multiple intermediate short videos according to the semantic similarity of the intermediate short videos to obtain a target short video.
[0045] That is, the process of combining multiple intermediate short videos according to the semantic similarity of the intermediate short videos to obtain a target short video can be regarded as a process of video content integration and creation, so as to generate a target short video with coherence and attractiveness:
[0046] In the video segmentation method provided in the embodiments of the present application, first, the initial video is preprocessed to generate multiple intermediate short videos, then the multiple intermediate short videos are subjected to semantic understanding processing, and finally, the multiple intermediate short videos are combined according to the semantic similarity of the intermediate short videos to obtain a target short video. The video segmentation method provided in the embodiments of the present application can achieve accurate and reasonable video scene segmentation in different situations such as different shot transitions, different scene changes, character subject changes, and severe occlusions, and avoid splitting the same action, the same character subject, and the same event into different video segments, thereby affecting the effect of video creation. Therefore, video scene segmentation can be realized in video creation, video editing, and similar theme video segment retrieval tasks, improving the processing efficiency and creativity of related tasks.
[0047] Further, in some embodiments of the present application, the combining the multiple intermediate short videos according to the semantic similarity of the intermediate short videos to obtain a target short video includes:
[0048] Combining adjacent intermediate short videos with related semantics according to the semantic similarity of the intermediate short videos to obtain a target short video.
[0049] Specifically, according to the semantic similarity of the intermediate short videos, adjacent intermediate short videos with related semantics can be combined to generate a target short video with coherence and integrity. The specific process is as follows: First, analyze the content of each intermediate short video and extract its key semantic information. This can be achieved through technologies such as natural language processing and computer vision to identify and understand elements such as dialogue, scene, and action in the video; then, calculate the semantic similarity between adjacent intermediate short videos. This can be achieved by comparing their key semantic information, such as theme, emotion, object, etc. Video segments with high similarity are more likely to belong to the same scene or storyline; finally, according to the calculation results of the semantic similarity, adjacent intermediate short videos with related semantics are combined. Through video editing technologies such as cutting and splicing, these segments are arranged in a certain order and logic to form a complete and coherent target short video. Such a target short video can better convey information, attract the attention of the audience, and achieve the expected communication effect.
[0050] Further, in some embodiments of the present application, combining adjacent intermediate short videos that are semantically related according to the semantic similarity of the intermediate short videos to obtain a target short video includes:
[0051] Calculating the embedding vector of the intermediate short video; and
[0052] Judging the semantic similarity of the intermediate short videos according to the embedding vector, and combining adjacent intermediate short videos that are semantically related to obtain a target short video.
[0053] Specifically, in order to judge the semantic similarity of intermediate short videos based on the embedding vector and combine adjacent videos that are semantically related to obtain a target short video, the following steps can be taken:
[0054] Step 1: Calculate the embedding vector of the intermediate short video
[0055] Feature extraction: For each intermediate short video, its key features need to be extracted. This can be achieved in various ways. For example, a deep learning model (such as a convolutional neural network CNN) can be used to extract the visual features of video frames, or natural language processing techniques can be used to extract the text features in the video (such as subtitles, voiceovers, etc.).
[0056] Embedding vector generation: Input the extracted features into an embedding model, which can map the high-dimensional features to a low-dimensional embedding space. The points (i.e., embedding vectors) in this embedding space can well represent the semantic information of the video. Commonly used embedding models include autoencoders, principal component analysis (PCA), etc., or more advanced cross-modal embedding models such as BERT and GPT for text and video.
[0057] Step 2: Judge the semantic similarity of the intermediate short videos
[0058] Similarity calculation: For each pair of adjacent intermediate short videos, calculate the similarity between their embedding vectors. This can be achieved through methods such as cosine similarity and Euclidean distance. Cosine similarity is a commonly used metric method, which calculates the cosine value of the angle between two vectors, and the closer the value is to 1, the more similar they are.
[0059] Similarity judgment: According to the result of the similarity calculation, set a threshold to judge which videos are semantically related. If the similarity of two videos exceeds this threshold, they are considered to be semantically related.
[0060] Step 3: Combine adjacent intermediate short videos that are semantically related
[0061] Video sorting and combination: According to the results of similarity judgment, semantically related adjacent intermediate short videos are arranged in order. This can be achieved through greedy algorithms, dynamic programming and other methods to ensure that the combined video is smooth and coherent.
[0062] Video editing and generation: Use video editing software or automated editing tools to edit and splice the arranged video clips to generate a complete target short video. In this process, you can also add transition effects, background music and other elements to enhance the viewing experience of the video.
[0063] In summary, by calculating the embedding vector of the intermediate short video and judging its semantic similarity, semantically related adjacent videos can be effectively combined to generate a target short video with coherence and integrity. This method has broad application prospects in the fields of video content creation and video recommendation. The code of the video scene segmentation algorithm based on semantic understanding is shown in the figure below. The video semantic understanding module
[0064] Its function is to perform semantic understanding on each short video clip generated by the transition detection module based on the video understanding model, convert the video information into an embedding vector representing the content of this video clip, and finally calculate the semantic similarity between each video clip according to the embedding vector. According to the video semantic similarity, semantically related adjacent video clips are combined to finally obtain reasonably segmented video clips.
[0065] Further, in some embodiments of the present application, judging the semantic similarity of the intermediate short videos according to the embedding vectors, and combining semantically related adjacent intermediate short videos to obtain the target short video includes:
[0066] When the embedding vector is greater than a preset threshold, it is determined that the adjacent intermediate short videos are semantically similar, and semantically related adjacent intermediate short videos are combined to obtain a target short video.
[0067] As mentioned above, for each pair of adjacent intermediate short videos, the similarity between their embedding vectors is calculated. This can be achieved through cosine similarity, Euclidean distance, etc. Cosine similarity is a commonly used measurement method, which calculates the cosine value of the angle between two vectors. The closer the value is to 1, the more similar they are.
[0068] According to the result of similarity calculation, a threshold is set to determine which videos are semantically related. If the similarity of two videos exceeds this threshold, they are considered to be semantically related, and the semantically related adjacent intermediate short videos are combined to obtain the target short video.
[0069] Furthermore, ifFigure 2 As shown, in some embodiments of the present application, the semantic understanding processing of the multiple intermediate short videos includes:
[0070] Performing semantic understanding processing on the multiple intermediate short videos through a video understanding module.
[0071] Specifically, input the multiple intermediate short videos into the video understanding module. For effective semantic understanding, it may be necessary to preprocess the video, such as video frame extraction, image enhancement, audio processing, etc., to improve the accuracy and efficiency of subsequent processing. The video understanding module extracts and represents features of the intermediate short videos, and performs semantic understanding and analysis, such as scene recognition, action recognition, object recognition and tracking, etc., so as to provide reference and basis for subsequent short video splicing.
[0072] Furthermore, in some embodiments of the present application, the preprocessing of the initial video to generate multiple intermediate short videos includes:
[0073] Performing transition detection on the initial video to generate multiple intermediate short videos.
[0074] Specifically, by performing transition detection on the initial video, we can achieve refined segmentation of the video. The specific process includes: First, extract video frames and perform feature extraction to obtain the visual differences between video frames; then, through methods such as mutation detection and gradual change detection, accurately identify the transition points in the video and further identify the transition types; finally, according to the transition detection results, segment the initial video into multiple intermediate short videos containing complete scenes, actions or events, and perform editing processing and number storage.
[0075] This process not only helps the understanding and analysis of video content, but also provides convenience for subsequent video editing, creation and recommendation. The intermediate short videos generated through transition detection can be used as the basis for video summary, highlight segment extraction, etc., to improve the efficiency and accuracy of video processing. At the same time, these intermediate short videos also provide a basis for the combination of video content.
[0076] Furthermore, in some embodiments of the present application, the performing transition detection on the initial video to generate multiple intermediate short videos includes:
[0077] Detecting the transition positions of the initial video on the time axis of the initial video; and
[0078] Dividing the initial video into the multiple intermediate short videos according to the transition boundaries of the transition positions of the initial video.
[0079] Specifically, in the initial video processing, the transition positions of the video are first detected on the timeline. This step involves a detailed analysis of the video content. By identifying transition features such as scene changes, shot transitions, or visual effect changes, the specific time points where transitions occur are determined. Subsequently, based on the detected transition positions, the transition boundaries are defined. These boundaries are crucial for video segmentation as they divide the video into different segments or scenes. Finally, based on these transition boundaries, the initial video is segmented into multiple intermediate short videos. Each intermediate short video independently contains a complete scene or event, providing a convenient material basis for subsequent video editing, analysis, or creation.
[0080] It should be noted that the video scene segmentation method shown in the embodiments of this application is a general one. The transition detection model and the video understanding model therein can be replaced with any relevant models. For example, the transition detection model can be replaced with video segmentation methods based on shots and actions, and the video understanding model can be replaced with any open-source or closed-source model.
[0081] Figure 3 is the structural block diagram of the video segmentation device 200 provided by the embodiments of this application. As Figure 3 shown, the video segmentation device 200 of the embodiments of this application includes: a preprocessing module 210, a semantic understanding module 220, and a short video generation module 230, where:
[0082] The preprocessing module 210 is used to preprocess the initial video to generate multiple intermediate short videos;
[0083] The semantic understanding module 220 is used to perform semantic understanding processing on the multiple intermediate short videos; and
[0084] The short video generation module 230 combines the multiple intermediate short videos according to the semantic similarity of the intermediate short videos to obtain the target short video.
[0085] In the video segmentation device provided by the embodiments of this application, first, the initial video is preprocessed to generate multiple intermediate short videos. Then, semantic understanding processing is performed on the multiple intermediate short videos. Finally, according to the semantic similarity of the intermediate short videos, the multiple intermediate short videos are combined to obtain the target short video. The video segmentation method provided by the embodiments of this application can achieve accurate and reasonable video scene segmentation in different situations such as different shot transitions, different scene changes, character subject changes, and severe occlusions, avoiding splitting the same action, the same character subject, and the same event into different video segments, thus affecting the effect of video creation. Therefore, video scene segmentation can be realized in video creation, video editing, and the task of retrieving video segments with similar themes, improving the processing efficiency and creativity of related tasks.
[0086] It should be noted that the specific implementation of the video segmentation device in the embodiments of this application is similar to that of the video segmentation method in the embodiments of this application. For details, please refer to the description in the method section and will not be elaborated here.
[0087] Figure 4 FIG. is a schematic structural diagram of an electronic device 300 according to an embodiment of this application.
[0088] As Figure 4 shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage section 302 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0089] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as required. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as required so that a computer program read from it can be installed into the storage section 308 as required.
[0090] Specifically, according to an embodiment of this application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of this application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 309 and / or installed from the removable medium 311. When the computer program is executed by a central processing unit (CPU) 301, the above functions defined in the electronic device of this application are executed.
[0091] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electronic device, apparatus, or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0092] In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction-executing electronic device, apparatus, or device. And in this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction-executing electronic device, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of a processing receiving device, method, and computer program product according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based electronic device that executes the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0094] The units or modules involved in the embodiments of the present application can be implemented in software or in hardware. The described units or modules can also be provided in a processor, and when the processor executes the program, a video segmentation method is implemented:
[0095] Preprocess the initial video to generate a plurality of intermediate short videos;
[0096] Perform semantic understanding processing on the plurality of intermediate short videos; and
[0097] Combine the plurality of intermediate short videos according to the semantic similarity of the intermediate short videos to obtain a target short video.
[0098] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the foregoing programs are executed by one or more processors, the video segmentation method described in the present application is implemented:
[0099] Preprocess the initial video to generate a plurality of intermediate short videos;
[0100] Perform semantic understanding processing on the plurality of intermediate short videos; and
[0101] Combine the plurality of intermediate short videos according to the semantic similarity of the intermediate short videos to obtain a target short video.
[0102] As another aspect, the present application further provides a computer program product, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer program product stores one or more programs, and when the foregoing programs are executed by one or more processors, the video segmentation method described in the present application is implemented:
[0103] Preprocess the initial video to generate a plurality of intermediate short videos;
[0104] Perform semantic understanding processing on the plurality of intermediate short videos; and
[0105] Combine the plurality of intermediate short videos according to the semantic similarity of the intermediate short videos to obtain a target short video.
[0106] The above are only the preferred embodiments of the present application, which do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the application concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A video segmentation method, characterized in that, Including the following steps: Preprocess the initial video to generate multiple intermediate short videos; Perform semantic understanding processing on the multiple intermediate short videos; And According to the semantic similarity of the intermediate short videos, combine the multiple intermediate short videos to obtain the target short video.
2. The video segmentation method according to claim 1, wherein The step of combining the multiple intermediate short videos according to the semantic similarity of the intermediate short videos to obtain the target short video includes: Combining adjacent intermediate short videos with related semantics according to the semantic similarity of the intermediate short videos to obtain the target short video.
3. The video segmentation method according to claim 1, wherein The step of combining adjacent intermediate short videos with related semantics according to the semantic similarity of the intermediate short videos to obtain the target short video includes: Calculating the embedding vector of the intermediate short time; and Judging the semantic similarity of the intermediate short videos according to the embedding vector, and combining adjacent intermediate short videos with related semantics to obtain the target short video.
4. The video segmentation method according to claim 3, wherein The step of judging the semantic similarity of the intermediate short videos according to the embedding vector, and combining adjacent intermediate short videos with related semantics to obtain the target short video includes: When the embedding vector is greater than the preset threshold, it is determined that adjacent intermediate short videos are semantically similar, and adjacent intermediate short videos with related semantics are combined to obtain the target short video.
5. The video segmentation method according to claim 1, characterized in that The step of performing semantic understanding processing on the multiple intermediate short videos includes: Through a video understanding module, perform semantic understanding processing on the multiple intermediate short videos.
6. The video segmentation method according to claim 1, wherein The step of preprocessing the initial video to generate multiple intermediate short videos includes: Generate multiple intermediate short videos by detecting transitions in the initial video.
7. The video segmentation method according to claim 6, wherein The step of generating multiple intermediate short videos by detecting transitions in the initial video includes: Detect the transition positions of the initial video on the time axis of the initial video; and Divide the initial video into the multiple intermediate short videos according to the transition boundaries of the transition positions of the initial video.
8. A video segmentation device, characterized in that, Including: A preprocessing module for preprocessing the initial video to generate multiple intermediate short videos; A semantic understanding module for performing semantic understanding processing on the multiple intermediate short videos; And A short video generation module that combines the multiple intermediate short videos according to the semantic similarity of the intermediate short videos to obtain the target short video.
9. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor is used to implement the video segmentation method according to any one of claims 1-7 when executing the program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is used to implement the video segmentation method according to any one of claims 1-7.