Short video automatic generation method based on multi-level label matching
By using a multi-level tag matching method, long videos are automatically processed into short video clips and the script is parsed, which solves the problems of low efficiency and unstable quality in short video production and achieves efficient and consistent short video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LAINENG (HANGZHOU) E-COMMERCE CO LTD
- Filing Date
- 2026-03-03
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies rely on manual operation in short video production, resulting in low efficiency and unstable quality. This makes it difficult to meet the demand for efficient and large-scale short video content production, and it is also difficult to guarantee production quality and consistency.
By establishing a multi-level tag matching method, long video samples are obtained and split into short video segments. Semantic tag sets are attached, the script to be processed is parsed and the matching degree is calculated. The short video segments are automatically matched and aggregated to generate a complete short video.
It has achieved automated and efficient short video production, ensuring the integrity of the video's narrative logic and the consistency of its quality, and solving the problems of low efficiency and unstable quality of manual processing.
Smart Images

Figure CN122120574A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video content processing technology, and in particular to a method for automatically generating short videos based on multi-level tag matching. Background Technology
[0002] With the rapid development of the mobile internet, short videos have become a core medium for information dissemination and commercial marketing. In many fields such as e-commerce, online education, and brand promotion, there is an urgent need for the production of high-quality, large-scale short video content. These videos typically require a pre-written script, which details the narrative logic, key information points, and emotional rhythm. The core of the production process lies in transforming the textual descriptions in the script into concrete video visuals. This involves accurately retrieving and selecting video clips from a vast library of video footage that semantically align with each part of the script, and finally splicing them together to create a complete video work.
[0003] Currently, the mainstream implementation method in the industry heavily relies on manual operation. Specifically, production staff must first thoroughly read and understand the overall structure and detailed meaning of the script; then, based on their personal experience and subjective judgment, they browse and filter through video footage libraries to find video materials that match the script in terms of visual content and contextual expression; finally, they manually edit, arrange, and composite the clips using video editing software. This human-centric workflow can still function when dealing with personalized, small-batch production needs, but its inherent limitations are also quite prominent.
[0004] First, it is extremely inefficient. The process of manually reading scripts, understanding semantics, and retrieving materials is time-consuming and cannot meet the current market's demand for "rapid iteration and large-scale" production of short video content. For example, in order to maximize the value of a live broadcast or interview, it is often necessary to quickly derive dozens or even hundreds of short video segments with different themes from a long video that lasts for several hours. Manual methods are difficult to accomplish in this scenario.
[0005] Secondly, production quality and consistency are difficult to guarantee. Due to the reliance on the subjective judgment of the production staff, different people, or even the same person at different times, may have different understandings of the same script. This directly leads to large fluctuations in the accuracy of the content expression and the consistency of style and tone of the generated videos, making it difficult to achieve standardized high-quality output.
[0006] Therefore, there is an urgent need in the existing technology to provide a method that can automatically and intelligently achieve accurate matching between scripts and video materials and efficiently generate high-quality short videos, so as to overcome the inefficiency and uneven quality caused by relying on manual operation. Summary of the Invention
[0007] Therefore, it is necessary to provide a method for automatically generating short videos based on multi-level tag matching, addressing the aforementioned shortcomings of existing technologies.
[0008] This application provides a method for automatically generating short videos based on multi-level tag matching, including: Multiple long video samples are obtained, each long video sample is split into multiple short video segments, and a set of semantic tags consisting of multiple semantic tags is attached to each short video segment; wherein, there is a semantic hierarchy among the semantic tags, and the multiple semantic tags attached to a single short video segment can describe the content of the short video segment from different dimensions. Store all short video clips into the video clip library; The script to be processed is obtained and parsed to obtain multiple script fragments and a set of semantic tags for each script fragment; Based on the semantic tag set of each script segment and the semantic tag set of each short video segment in the video segment library, the matching degree between each script segment and each short video segment is calculated, and one or more target short video segments are matched from the video segment library according to the matching degree. All matched target short video segments are aggregated according to the semantic coherence of the script segments to generate a complete short video.
[0009] This application relates to a method for automatically generating short videos based on multi-level tag matching. By establishing a hierarchical semantic tagging system, it provides a unified, machine-understandable structured description for both video and script content. By breaking down long videos into segments with semantic tags, unstructured video data is transformed into searchable structured data. Through isomorphic parsing and tagging of the script, a mapping between text and video is achieved within the same semantic space. The introduction of a semantic tag-based matching mechanism replaces the manual understanding and searching process, automating the matching of materials. Aggregation according to the semantic order of the script ensures the completeness of the narrative logic in the generated video. The multi-level tag matching-based method for automatically generating short videos provided in this application solves the core problems of low efficiency and unstable quality associated with manual processing. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating a method for automatically generating short videos based on multi-level tag matching, as provided in an embodiment of this application. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0012] This application provides a method for automatically generating short videos based on multi-level tag matching.
[0013] Furthermore, the short video automatic generation method based on multi-level tag matching provided in this application does not limit the executing entity. Optionally, the executing entity of the short video automatic generation method based on multi-level tag matching provided in this application can be a short video automatic generation terminal. Specifically, the executing entity of the short video automatic generation method based on multi-level tag matching provided in this application can be one or more processors in the short video automatic generation terminal.
[0014] like Figure 1 As shown, in one embodiment of this application, the short video automatic generation method based on multi-level tag matching includes: S100: Obtain multiple long video samples, split each long video sample into multiple short video segments, and attach a semantic tag set consisting of multiple semantic tags to each short video segment. These semantic tags have a hierarchical semantic relationship, and the multiple semantic tags attached to a single short video segment can describe the content of the short video segment from different dimensions.
[0015] S200 stores all short video clips into the video clip library.
[0016] S300: Obtain and parse the script to be processed, and obtain multiple script fragments and a set of semantic tags for each script fragment.
[0017] S400: Based on the semantic tag set of each script segment and the semantic tag set of each short video segment in the video segment library, calculate the matching degree between each script segment and each short video segment, and match one or more target short video segments from the video segment library for each script segment according to the matching degree.
[0018] S500: All matched target short video segments are aggregated according to the semantic coherence order of the script segments to generate a complete short video.
[0019] Specifically, steps S100 to S200 involve constructing a video clip library. First, multiple long videos from live streams or explanations are acquired as samples. For each long video sample, it is then divided into multiple short video clips. For example, one short video clip might be a "swatch demonstration on an arm," and another might be a "lipstick ingredient explanation."
[0020] Next, a semantic tag set is generated for each short video segment. For example, for a "swatch demonstration" segment, its semantic tag set might include the following semantic tags: L1: Makeup, L2: Product Showcase, L3: Swatch. These semantic tags describe the same short video segment from different dimensions such as product type, content function, and specific actions, forming a hierarchical relationship, such as Makeup > Product Showcase > Swatch, refining and specifying layer by layer. Finally, all processed short videos are stored in a video segment library. The semantic tag set can be stored separately using a vector database. Semantic tags are represented in vector form in the vector database, with each semantic tag corresponding to a semantic vector. The hierarchical relationship between multi-level semantic tags is expressed through clustering relationships, vector distance, or parent-child tag mapping relationships in the vector space. In this example, the semantic vectors themselves, as well as the hierarchical relationships, are stored in the vector database to support efficient retrieval and hierarchical matching based on vector distance.
[0021] S300 is the script processing procedure. It receives a script to be processed. For example, scripts are generally complex and contain a lot of content; for brevity, a simple example will be used here. For instance, the script to be processed might be: "This lipstick has a silky smooth texture and doesn't dry out the skin. Our classic true red is very flattering." The file format of the script to be processed has already been pre-labeled with multi-level semantic tags; that is, the file format of the script to be processed has already been pre-attached with a set of semantic tags consisting of multiple semantic tags. The semantic tag set of the script to be processed and the semantic tag set of the short video have a isomorphic relationship, meaning they share the same generation principle and logic. This allows them to understand each other in the subsequent matching steps.
[0022] The file format of the script to be processed can be: { "script_content": "Not wearing lip gloss in autumn is a huge waste! Shade R300 is practically the color of a pristine white moonlight from a novel!", "script_split": [ { "content": "Not wearing lip gloss in autumn is a huge waste!" "tags": [ "Scene Anchoring Tags" ] }, { "content": "Shade R300 is practically the color of the white moonlight from a novel!" "tags": [ "Product Lens Label" ] } ]} The S300 can parse the script to be processed, obtaining the text content of the script and the total semantic tag set of the script. A Large Language Model (LLM) can be used to semantically segment the text content of the script, splitting it into two script segments, such as "This lipstick has a silky smooth texture and is not drying" and "Our classic true red is very flattering." Similarly, the total semantic tag set of the script is decomposed to obtain the semantic tag set of each script segment. The semantic tag set of the first script segment includes the following semantic tags: L1: Makeup, L2: Product description, L3: Texture description. The semantic tag set of the second script segment includes the following semantic tags: W1: Makeup, W2: Product description, W3: Shade description.
[0023] S400 is a process of pairing script snippets with short video snippets one by one. It requires comparing the semantic tag set of each script snippet with the semantic tag set of each short video snippet in the video snippet library. By calculating the matching degree, the most suitable short video snippet is found for each script snippet. For example, the script snippet "Our classic true red, very flattering" matches a short video snippet in the video snippet library containing three semantic tags: L1: Makeup, L2: Product Display, and L3: Color Swatch, with the content being a close-up of a model's lips. It's important to note that a script snippet may match only one short video snippet, or it may match multiple short video snippets. This is because a single script snippet might be too long and contain too much information, requiring multiple short video snippets to fully convey its information.
[0024] The final S500 process combines all matched target short video clips strictly according to the original semantic coherence of the script clips, generating a complete short video with coherent content and a high degree of consistency between visuals and text. The semantic coherence order refers to the narrative logic of the script itself.
[0025] In this embodiment, a hierarchical semantic tagging system is established to provide a unified, machine-understandable structured description for both video and script content. By breaking down long videos into segments with semantic tags, unstructured video data is transformed into searchable structured data. Through isomorphic parsing and tagging of the script, a mapping between text and video is achieved within the same semantic space. The introduction of a semantic tag-based matching mechanism replaces the manual understanding and searching process, automating material matching. Aggregation according to the semantic order of the script ensures the integrity of the narrative logic in the generated video. The short video automatic generation method based on multi-level tag matching provided in this application solves the core problems of low efficiency and unstable quality associated with manual processing.
[0026] In one embodiment of this application, S100 includes obtaining multiple long video samples, splitting each long video sample into multiple short video segments, and attaching a semantic tag set consisting of multiple semantic tags to each short video segment, including: S110, select a long video sample.
[0027] S120, determine whether the long video sample carries valid audio data.
[0028] S130, if the long video sample carries valid audio data, then the speech is separated from the long video sample. Through speech recognition, speech endpoint detection and semantic coherence analysis of the large language model, the semantically complete speech unit and its accurate time boundary are determined, thereby obtaining multiple short video segments after splitting.
[0029] S140, if the long video sample does not carry valid audio data, the scene transition point is detected by analyzing the changes in visual features between consecutive video frames, thereby obtaining multiple short video segments after splitting.
[0030] S150, return to S100, that is, return to the step of selecting a long video sample until each long video sample is split into multiple short video segments.
[0031] Specifically, S110 can select a sample sequentially from the queue of long video samples to be processed, such as a long makeup tutorial video. Alternatively, it can select a sample randomly.
[0032] In S120, tools such as FFmpeg can be used to check whether the long video sample file contains a valid audio track and whether the audio contains human voice. Optionally, this can be determined by calculating the audio amplitude or using a simple silence detection. If the long video sample file contains a valid audio track and the audio contains human voice, then the long video sample is confirmed to carry valid audio data.
[0033] If the long video sample carries valid audio data, it proceeds to the first decomposition process. This involves separating the audio from the long video sample, using an ASR (Acoustic Recognition) tool to convert the speech into text, and then using a Voice Endpoint Detection (VAD) tool to initially determine the start and end points of speech segments. Subsequently, an LLM (Large Language Model) is used to perform semantic analysis on the text, merging semantically related short sentences, and finally determining the precise time boundaries before video cutting.
[0034] If a long video sample does not carry valid audio data (such as a video with only background music or a silent, fast-cut video), it proceeds to the second decomposition process. This involves using libraries such as OpenCV to extract video frames, calculating inter-frame differences, detecting shot transition points by setting a threshold, and cutting the video accordingly.
[0035] After completing the decomposition of a long video sample, return to S110 to select the next long video sample, until all long video samples have been processed.
[0036] Of course, a long video sample will inevitably have both parts carrying audio and parts not carrying audio. Therefore, a combination of S130 and S140 can be used to process the long video sample. That is, the part carrying valid audio data is decomposed using S130, and the part not carrying valid audio data is decomposed using S140. It is possible that a long video sample will eventually be decomposed into 50 short video segments, of which 30 short video segments carry valid audio data and 20 short video segments do not carry valid audio data.
[0037] In this embodiment, differentiated processing of different video types is achieved by determining whether long videos carry valid audio data. Videos with audio are processed using a speech analysis path, while videos without audio are processed using a visual analysis path, ensuring that each type receives the most suitable processing method. This adaptive processing mechanism improves the accuracy and completeness of video decomposition and avoids the performance loss that occurs when using a single method to process mixed-type videos. The loop processing mechanism supports batch processing of multiple long video samples, improving the efficiency of building the video clip library.
[0038] In one embodiment of this application, S130 includes, namely, separating speech from long video samples, determining semantically complete speech units and their accurate temporal boundaries through speech recognition, speech endpoint detection, and semantic coherence analysis of a large language model, thereby obtaining multiple short video segments after segmentation, including: S131, the audio file is separated from the long video sample.
[0039] S132, perform speech recognition and speech endpoint detection on the speech file to obtain the initial speech text segment and its corresponding timestamp.
[0040] S133, The initial speech text segment is input into the large language model for semantic coherence analysis.
[0041] S134, based on the results of semantic coherence analysis, semantically related continuous speech text segments are merged into semantically complete speech units, and the accurate start and end timestamps of the semantically complete speech units in the long video sample are determined.
[0042] S135, based on the accurate start and end timestamps, the long video sample is split into corresponding short video segments.
[0043] Specifically, this embodiment describes in detail the disassembly process of a long video sample carrying valid audio data.
[0044] S131 is the process of separating audio, which can be done using FFmpeg to extract individual audio files (such as WAV format) from a long video.
[0045] The following two tools are mainly used in speech recognition and endpoint detection in S132: Automatic Speech Recognition (ASR): When used, an audio file is input into the ASR engine. It works by converting the audio signal into text using acoustic and language models, and outputting text segments with timestamps. For example, it recognizes "this lipstick" [00:01:00-00:01:04], "very textured" [00:01:04-00:01:07], and "silky smooth" [00:01:07-00:01:10]. 00:01:00 refers to the timestamp of 00:01:00. [00:01:00-00:01:04] refers to the time period from 00:01:00 to 00:01:04.
[0046] Speech Endpoint Detection (VAD): This technology is used to detect intervals between speech and non-speech (silence or noise) in audio. By analyzing features such as short-time energy and zero-crossing rate, it can more accurately locate the start and end of speech activity, which helps to correct minor deviations in ASR timestamps.
[0047] In summary, speech recognition and endpoint detection can convert audio signals into text. The final text is a text with timestamps and the speaker's name.
[0048] For example, in the text “2 spk0 00:00:01,670 --> 00:00:03,270 This bottle is no more than fifty yuan”, “2 spk0” refers to the first person’s words in the second line (2) (spk0, i.e., speak0), “00:00:01,670 --> 00:00:03,270” is the timestamp, and “This bottle is no more than fifty yuan” is the text content.
[0049] S134 is the process of semantic coherence analysis. In this step, text fragments that are sequential in time and have short intervals (such as the three short sentences "this lipstick", "very textured", and "silky smooth" in the example above) generated by ASR and VAD are input into the LLM model. Based on its understanding of language structure and context, the LLM model determines whether these fragments semantically belong to a complete semantic group. How the LLM model achieves semantic understanding is a common-sense technique. Those skilled in the art do not need a detailed explanation of how the LLM model achieves semantic understanding to implement the technical content mentioned in this application without difficulty. Furthermore, how the LLM model achieves semantic understanding is not the focus of this application. Therefore, a detailed explanation of how the LLM model achieves semantic understanding is not provided here.
[0050] Optionally, after semantic coherence analysis, the LLM model outputs multiple consecutive semantically related speech-text segments. The first method for determining the breakpoints between different batches is undoubtedly an understanding of the language structure and context. The second method is to identify a text segment containing a specific ending phrase as the endpoint of that batch of consecutive semantically related speech-text segments, using this endpoint as the breakpoint to segment different batches. For example, when a text segment containing the phrase "123 link up" is identified, it represents the official end of multiple consecutive product promotion speech-text segments, and the conversation may transition to the next batch of multiple consecutive price introduction speech-text segments.
[0051] S135 is the step of merging and determining time boundaries. After LLM determines that "this lipstick has a very silky texture" is a complete semantic unit, the earliest time (00:01:00) and the latest time (00:01:10) among the start timestamps of all the short sentences in the unit are used as the accurate time boundaries of this semantic unit.
[0052] Finally, based on this merged time boundary (00:01:00-00:01:10), the corresponding segments in the long video are cut out using FFmpeg commands to generate an independent short video segment.
[0053] In this embodiment, key issues in video-with-audio decomposition are addressed by combining speech recognition, speech endpoint detection, and large language model semantic analysis. Speech recognition converts audio into text, speech endpoint detection determines speech segment boundaries, and the large language model analyzes semantic coherence. The collaborative work of these three technologies ensures that the decomposition results maintain both technical accuracy and semantic integrity. This method overcomes the semantic fragmentation problem caused by relying solely on speech recognition and avoids the semantic breaks that may occur with endpoint detection alone, providing short video clips with complete semantic expression for subsequent processing.
[0054] In one embodiment of this application, S134 includes, based on the result of the semantic coherence analysis, merging semantically related continuous speech text segments into semantically complete speech units, and determining the accurate start and end timestamps of the semantically complete speech units in the long video sample, including: S134a: Obtain multiple consecutive semantically related speech-text segments output by the large language model.
[0055] S134b, obtain the time interval between every two adjacent speech text segments in a series of semantically related speech text segments.
[0056] S134c, when the time interval between any two adjacent speech text segments is less than a preset duration, multiple consecutive semantically related speech text segments are merged into one speech unit.
[0057] S134d, treats the voice unit as a short video clip.
[0058] Specifically, this embodiment further introduces the specific rules for generating semantic units using the LLM model.
[0059] The result of the semantic coherence analysis in S133 is a series of consecutive semantically related speech-text segments. Subsequently, S134a to S134d to S135 are performed on each series of consecutive semantically related speech-text segments.
[0060] Following the above embodiments, the LLM model has determined that the three consecutive speech text segments A "this lipstick" (00:01:00-00:01:04), B "very textured" (00:01:04-00:01:07), and C "silky smooth" (00:01:07-00:01:10) are semantically related.
[0061] Next, the time interval between adjacent segments will be calculated. For example, the interval between speech text segment B and speech text segment A is 0 seconds (immediately adjacent), and the interval between speech text segment C and speech text segment B is 0 seconds (immediately adjacent).
[0062] The system presets a maximum allowed interval threshold, such as 2 seconds.
[0063] Since the time interval (0 seconds) of all adjacent segments is lower than the preset maximum allowable interval threshold (2 seconds), and the semantics are determined to be related by the LLM model, the system merges the three speech text segments A, B, and C into a complete speech unit "This lipstick has a very silky texture", with a time boundary of 00:01:00 to 00:01:10.
[0064] The merged speech unit will be treated as an independent short video segment for subsequent processing. If a speech-text segment is more than 2 seconds apart from the previous speech-text segment, the system will not merge them, even if the LLM model considers them semantically related, to ensure that the duration of each speech-text segment is kept within a reasonable range.
[0065] After the LLM model processes the three speech-text segments A, B, and C, it continues to process the next batch of multiple consecutive semantically related speech-text segments in sequence until the initial speech-text segments have been completely processed.
[0066] In this embodiment, intelligent merging of speech units is achieved through a dual mechanism of time interval judgment and semantic relevance analysis. The preset duration threshold distinguishes between normal speech pauses and semantic transition boundaries, ensuring that the merged speech units maintain both temporal continuity and semantic integrity. Compared to fixed-duration segmentation methods, the short video clips generated by this method better conform to the rules of natural language expression, avoid semantic breaks caused by mechanical segmentation, and improve the usability of the clips as independent semantic units.
[0067] In one embodiment of this application, S140 includes, namely, detecting scene transition points by analyzing visual feature changes between consecutive video frames, thereby obtaining multiple short video segments after splitting, including: S141, Extract the continuous video frame sequence of the long video sample.
[0068] S142, calculate the visual feature difference between any two consecutive adjacent video frames.
[0069] S143, when the visual feature difference between two adjacent consecutive video frames is greater than or equal to the preset visual feature difference threshold, the timestamp of the previous video frame is used as the scene transition point.
[0070] S144, based on all detected scene transition points, the long video sample is split into multiple scene segments.
[0071] S145 treats each scene segment as a short video clip.
[0072] Specifically, this embodiment describes in detail the disassembly process of long video samples that do not carry valid audio data.
[0073] In S141, FFmpeg can be used to extract image sequences from long, silent videos at a fixed frame rate (e.g., 10 frames per second). Each image is a video frame.
[0074] In S142, for two adjacent consecutive video frames, OpenCV can be used to calculate their color histograms, and the Bach distance or correlation coefficient between the two color histograms can be calculated as the visual feature difference. The greater the difference, the more drastic the change in the content of the two frames. Of course, this is just one implementation method; any form of visual feature difference algorithm can be used without limitation.
[0075] Next, S143 describes the process of detecting scene transition points. A preset visual feature difference threshold can be set (e.g., 0.6, which is an empirical value that can be obtained through experiments or questionnaires). When the calculated visual feature difference between two adjacent consecutive video frames is greater than or equal to 0.6, a scene transition is determined to have occurred, and the timestamp of the previous video frame is recorded as the scene transition point. For example, at timestamp 00:02:15, the scene suddenly switches from the overall appearance of the lipstick to a close-up of the lips, and the calculated visual feature difference reaches 0.8. Therefore, this time point 00:02:15 is recorded as a scene transition point.
[0076] S144 is the specific operation process for splitting the video into segments. In S144, all detected scene transition points are collected, along with the start and end times of the long video sample. FFmpeg is then used to cut the long video into multiple scene segments, with the scene transition point being the cut-off point. Each scene segment represents a visually coherent shot.
[0077] In this embodiment, scene transition point detection in audio-visual videos is achieved by analyzing visual feature changes between consecutive video frames. Calculating the visual feature differences between adjacent frames allows for the objective identification of moments when significant changes occur in the video content. The preset difference threshold balances the sensitivity and specificity of the detection, accurately capturing real-world scene changes while avoiding misjudgments caused by minor image fluctuations. This method provides an effective decomposition scheme for videos primarily containing visual content, thus expanding the coverage of the video segment library.
[0078] In one embodiment of this application, S100 further includes, namely, acquiring multiple long video samples, splitting each long video sample into multiple short video segments, and attaching a semantic tag set consisting of multiple semantic tags to each short video segment, and further includes: S161, Select a short video clip and create a semantic tag set corresponding to the short video clip.
[0079] S162, identify the main content in the short video segment through a visual analysis model to generate at least one semantic tag, and incorporate the at least one semantic tag into the semantic tag set corresponding to the short video segment.
[0080] S163, Analyze the text content corresponding to the short video segment using a text analysis model to generate at least one semantic tag, and incorporate the at least one semantic tag into the semantic tag set corresponding to the short video segment.
[0081] The generated semantic tags form a hierarchical or parallel relationship with other existing semantic tags in the semantic tag set.
[0082] Specifically, this embodiment describes the process of generating semantic tags using multimodal methods.
[0083] S161 is the process of selecting and initializing short video clips. A pre-decomposed short video clip, such as a 5-second "arm color swatch" clip, is selected, and an empty semantic tag set is generated for it.
[0084] S162 generates semantic tags corresponding to the main content through visual analysis. Several possible methods are available and no restrictions are imposed. Optionally, keyframes are extracted from the short video clip and input into a pre-trained image classification model (such as ResNet) or object detection model (such as YOLO). The model identifies the main objects in the scene as "lipstick" and "arm," and the action as "applying." The model outputs a high-confidence category, such as "lipstick." In this step, visually relevant tags are generated and added to the short video clip based on a predefined tag hierarchy mapping, such as L1: makeup, L2: product display, and L3: color swatch.
[0085] S163 generates semantic tags corresponding to text content through text analysis. For segments with speech, it uses ASR-transcribed text. For segments without speech, it uses video OCR to recognize screen text. The obtained text is then input into an NLP model (such as the BERT model) for keyword extraction or text classification. For example, after analyzing the text "This velvet texture is really high-end," the keyword "velvet texture" is extracted. Based on this, the system generates and adds text-related semantic tags, such as L1: beauty, L2: product description, and L3: texture description.
[0086] S160 may also include S164, constructing a semantic tag set. At this point, the semantic tag set for the short video clip contains multiple semantic tags from both visual and textual modalities. Among these semantic tags, L1: Beauty, L2: Product Display, and L2: Product Narration form a hierarchical relationship. L2: Product Display and L2: Product Narration, both being secondary tags, can be viewed as parallel relationships describing the clip from different dimensions.
[0087] In this embodiment, a comprehensive understanding of short video clip content is achieved through multimodal processing combining visual and textual analysis. Visual analysis identifies the main content in the image, while textual analysis parses the corresponding linguistic information. The information from these two sources complements each other, constructing a multi-faceted description of the video clip. The generated semantic tags form hierarchical or parallel relationships, establishing a structured representation of the clip content. This multimodal tag generation mechanism provides the necessary information foundation for subsequent high-precision matching.
[0088] In one embodiment of this application, S400 includes calculating the matching degree between each script segment and each short video segment based on the semantic tag set of each script segment and the semantic tag set of each short video segment in the video segment library, and matching one or more target short video segments from the video segment library for each script segment according to the matching degree, including: S411, retrieve a script fragment and a short video fragment to be matched.
[0089] S412, convert the semantic tag set of the script fragment and the semantic tag set to be matched into corresponding feature vectors respectively.
[0090] S413, calculate the duration matching factor by combining the estimated duration of the script segment with the actual duration of the short video segment.
[0091] S414, Based on the feature vector and the duration matching factor, a weighted calculation is performed to obtain the comprehensive matching score of the short video segment.
[0092] Specifically, this embodiment describes one implementation of the matching degree calculation method.
[0093] S411 retrieves a one-to-one match. For example, a script snippet S (containing the semantic content: "This lipstick is very moisturizing", estimated duration 3 seconds) and a short video snippet V.
[0094] S412 can use word embedding models (such as Word2Vec or BERT) to convert each semantic tag (such as "beauty", "product description", "moisturizing") of script fragment S and video fragment V into a high-dimensional numerical vector (e.g., a 300-dimensional vector), thus vectorizing the semantic tags in both fragments.
[0095] The weighted average of all the tag vectors of the same segment can be calculated (the weight can be determined by the tag level, with higher levels having higher weights) to obtain the feature vector Vec_s of script segment S and the feature vector Vec_v of short video segment V.
[0096] In S413, the duration matching factor can be calculated using Formula 1.
[0097] T=1-|D_s-D_v| / (D_s+D_v) Formula 1.
[0098] Where T is the duration matching factor. D_s is the estimated duration of the script segment (e.g., 3 seconds). D_v is the actual duration of the short video segment (e.g., 3.5 seconds). The calculation yields T = 1 - |3 - 3.5| / (3 + 3.5) ≈ 0.92. The closer T is to 1, the better the durations of the short video segment and the script segment match. The greater the difference in duration between the short video segment and the script segment, the closer the T value is to 0.
[0099] The core significance of the T parameter lies in ensuring that the matched short video clips are appropriate in terms of "duration," avoiding awkward situations such as the screen switching before the voiceover ends or the screen still spinning after the voiceover ends. It ensures the smoothness of the matching algorithm's rhythm.
[0100] In S414, cosine similarity can be used to calculate the similarity of semantic vectors, and then a weighted fusion method can be used to calculate the comprehensive matching score.
[0101] Use Formula 2 to calculate the similarity of semantic vectors.
[0102] Sem_Sim=cosine_similarity(Vec_s,Vec_v) Formula 2.
[0103] Where Sem_Sim represents the semantic similarity. Vec_s is the feature vector of the script segment. Vec_v is the feature vector of the short video segment. cosine_similarity is the notation for calculating cosine similarity. Cosine similarity focuses on the difference in direction between two vectors, without being affected by their length (magnitude), making it very suitable for measuring semantic similarity.
[0104] The core significance of the Sem_Sim parameter lies in ensuring that the matched video footage and script text are consistent in meaning, which is the foundation for the accuracy of the matching algorithm.
[0105] Use Formula 3 to calculate the overall matching score.
[0106] W=W1×Sem_Sim+W2×T Formula 3.
[0107] Where W is the overall matching score, and W1 and W2 are preset weights. W1 is the weight of semantic vector similarity, and W2 is the weight of the duration matching factor. For example, W1=0.8 and W2=0.2 indicates a greater emphasis on semantic matching. W1 and W2 are user-defined.
[0108] In this embodiment, a quantitative matching evaluation method is established by converting semantic tags into feature vectors and introducing a duration matching factor. Feature vector conversion makes semantic similarity calculation more accurate, while the duration matching factor ensures the consistency of content across the time dimension. The weighted calculation integrates semantic and duration features, and the resulting comprehensive matching score fully reflects the degree of matching between the script segment and the short video segment. This quantitative evaluation method provides an objective basis for material selection and reduces the subjectivity of the matching process.
[0109] In one embodiment of this application, S400 further includes calculating the matching degree between each script segment and each short video segment based on the semantic tag set of each script segment and the semantic tag set of each short video segment in the video segment library, and matching one or more target short video segments from the video segment library for each script segment according to the matching degree, and further includes: S420: Select a script segment and obtain the target number and score threshold of the corresponding short video segment.
[0110] S430, determine whether the number of priority matching segments is greater than or equal to the target number of short video segments.
[0111] S441, if the number of segments to be matched first is greater than or equal to the number of target short video segments, then select the N short video segments with the highest overall matching score as the target short video segments to match this script segment. N is a positive integer and N is equal to the number of target short video segments.
[0112] S442, if the number of priority matching segments is less than the number of short video segment targets, then supplementary matching is performed from several short video segments that are the same as the script segment in terms of semantic tags at a lower level, and the priority matching segments and the supplementary matching short video segments are used together as the target short video segments that match the script segment.
[0113] Specifically, this embodiment describes the execution logic of the matching strategy.
[0114] S420 is the process of setting matching targets. For example, for a script segment, the number of short video segments that it needs to match is preset to N (e.g., N=1) and a minimum acceptable overall matching score threshold is set (e.g., 0.7).
[0115] The following describes the priority matching process, namely S430, S441, and S442.
[0116] First, all video clips that are completely identical to the script clips in higher-level tags (e.g., L1 and L2) are selected and placed into a priority matching pool. Their overall match score is then calculated. Assuming the script clip tags are L1: Beauty, L2: Product Description, and L3: Efficacy, then the short video clips in the priority matching pool must also have L1: Beauty and L2: Product Description.
[0117] If the priority matching pool meets the requirements: if the number of short video segments with a comprehensive matching degree of more than 0.7 in the priority matching pool is greater than or equal to N (1 in this example), then the N highest-scoring segments are selected as the target short video segments in descending order of their scores.
[0118] If the priority matching pool does not meet the requirements: If there are no short video clips with a comprehensive matching degree exceeding 0.7 in the priority matching pool, or if the number of short video clips with a comprehensive matching degree exceeding 0.7 is less than N, then the supplementary matching mechanism will be activated. Supplementary matching will relax the conditions, requiring only that the video clip and the script clip have the same L1 tag (i.e., L1: beauty), even if the L2 or L3 tags are different. Clips with high comprehensive matching degrees will be selected from this pool to supplement the script clip, and the short video clips from the priority matching pool and the supplementary matching pool will be used together as the target short video clip for matching the script clip.
[0119] Finally, all target short video segments matched by the script fragment (which may come from the priority match or may include supplementary matches) are output.
[0120] In this embodiment, a combination of priority matching and supplementary matching balances the requirements for matching quality and quantity. Priority matching ensures the quality of core content, while supplementary matching provides additional materials when necessary. This dynamically adjusted matching strategy adapts to different needs. The target quantity requirement ensures the structural integrity of the generated video, and the hierarchical matching mechanism demonstrates good adaptability when dealing with limited material libraries or uneven content distribution.
[0121] In one embodiment of this application, the segment to be matched first is a short video segment that is exactly the same as the script segment in terms of semantic tags at a higher level.
[0122] Specifically, in this embodiment, the criteria for priority matching in the aforementioned embodiments are further clarified.
[0123] Using the previous example, the script segment tags are L1: Beauty, L2: Product Description, and L3: Efficacy.
[0124] Higher-level semantic tags refer to tags that are closer to the root node in the hierarchy tree, are more abstract, and determine the major categories of content, namely L1 and L2.
[0125] "Completely identical" means that the video clip must have both the L1: Beauty and L2: Product Description tags. Even if it has a very relevant L3 tag (such as L3: Texture), it will not be included in the priority match pool if its L2 tag is product demonstration rather than product description.
[0126] This standard ensures that the system first seeks video footage that is highly consistent with the script's intent in terms of content category and usage scenario.
[0127] Therefore, this embodiment embodies the strategy's degradation path: we always prioritize using abstract and high-level tags for matching, but when a perfect match cannot be found, we expand the search scope by reducing the total number of levels that need to be matched (i.e., abandoning the requirement for lower-level, more specific tags). Reducing the total number of levels that need to be matched is equivalent to the meaning of "fewer levels of semantic tags" in the supplementary matching mechanism.
[0128] In this embodiment, by explicitly defining the priority matching criterion as "identical higher-level semantic tags," a quality-first principle is established in the matching process. Higher-level tags correspond to the main category and core features of the content; consistency in these tags ensures that the matched segment and the script align in terms of theme and main content. This matching criterion grasps the key elements of content matching, providing a fundamental guarantee for the overall quality of the generated video.
[0129] In one embodiment of this application, S500 includes aggregating all matched target short video segments according to the semantic coherence order of the script segments to generate a complete short video, comprising: S510, with smoothness optimization.
[0130] The smoothness optimization process includes at least one or more of the following: adding transition effects between adjacent target short video segments, aligning the scene transition points of the target short video segments with the beat points of the background music, and applying fade-in and fade-out effects to the audio track of the target short video segments.
[0131] Specifically, this embodiment defines quality optimization measures for the aggregation stage.
[0132] The following three smoothness optimization techniques are all available: Transition effects: When the system detects that two adjacent video clips have significant differences in visual content (for example, by calculating the histogram difference between the first and last frames), in order to avoid abrupt transitions, it will automatically add transition effects such as "fade in / fade out", "dissolve", or "slide" between the two clips to make the transition more natural.
[0133] Rhythm Matching: If the final video includes background music, the system uses audio analysis libraries such as Aubio to detect the beats of the BGM. Then, it tries to align the scene transitions of short video clips with these beats, thereby enhancing the video's rhythm and visual appeal.
[0134] Audio Fade In / Out: Apply a brief volume transition effect (e.g., a 0.3-second fade in and fade out) at the beginning and end of the audio track of each video clip. This eliminates abrupt start or stop in audio caused by clip splicing, ensuring auditory continuity.
[0135] In this embodiment, smoothness optimization processing is introduced during the compositing stage to improve the visual quality of the automatically generated video. Transition effects eliminate visual jumps between segments, alignment of switching points with music beats enhances the sense of audiovisual synchronization, and audio fade-in / fade-out ensures natural sound transitions. These processes improve the final video output from both technical and experiential perspectives, enabling the automatically generated video to achieve a visual quality approaching that of professionally produced videos.
[0136] Before the aggregation step, i.e. before S500, the following steps are also included: S450 performs standardized processing on all matched target short video clips to ensure consistency in technical specifications.
[0137] Optionally, S450 may include video normalization. Specifically, FFmpeg is used to normalize the resolution of all segments to the target specification (e.g., 1440x2560), the frame rate to 60fps, and the video bitrate to 16000kbps.
[0138] Optionally, the S450 can include audio normalization: specifically, using an audio processing library (such as SoX) to unify the audio loudness of all segments to -14 LUFS, limit the true peak to below -1 dBTP, unify the audio bitrate to 192 kbps, and unify the sampling rate to 48 kHz.
[0139] This step can effectively solve problems such as choppy video playback and fluctuating sound volume caused by different material sources.
[0140] The technical features of the above embodiments can be combined arbitrarily, and the execution order of the method steps is not restricted. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0141] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for automatically generating short videos based on multi-level tag matching, characterized in that, include: Multiple long video samples are obtained, each long video sample is split into multiple short video segments, and a set of semantic tags consisting of multiple semantic tags is attached to each short video segment; wherein, there is a semantic hierarchy among the semantic tags, and the multiple semantic tags attached to a single short video segment can describe the content of the short video segment from different dimensions. Store all short video clips into the video clip library; The script to be processed is obtained and parsed to obtain multiple script fragments and a set of semantic tags for each script fragment; Based on the semantic tag set of each script segment and the semantic tag set of each short video segment in the video segment library, the matching degree between each script segment and each short video segment is calculated, and one or more target short video segments are matched from the video segment library according to the matching degree. All matched target short video segments are aggregated according to the semantic coherence of the script segments to generate a complete short video.
2. The method for automatically generating short videos based on multi-level tag matching according to claim 1, characterized in that, The process of acquiring multiple long video samples, splitting each long video sample into multiple short video segments, and attaching a semantic tag set consisting of multiple semantic tags to each short video segment includes: Select a long video sample; Determine whether the long video sample carries valid audio data; If the long video sample carries valid audio data, the speech is separated from the long video sample. Through speech recognition, speech endpoint detection and semantic coherence analysis of the large language model, the semantically complete speech unit and its accurate time boundary are determined, and then multiple short video segments are obtained after splitting. If the long video sample does not carry valid audio data, the scene transition points are detected by analyzing the changes in visual features between consecutive video frames, thereby obtaining multiple short video segments after splitting. Return to the previous step and select a long video sample until each long video sample is split into multiple short video segments.
3. The method for automatically generating short videos based on multi-level tag matching according to claim 2, characterized in that, The process involves separating speech from long video samples, and through speech recognition, speech endpoint detection, and semantic coherence analysis using a large language model, determining semantically complete speech units and their accurate temporal boundaries, thereby obtaining multiple short video segments after segmentation, including: The audio file was extracted from the long video sample; The speech file is subjected to speech recognition and speech endpoint detection to obtain the initial speech text segment and its corresponding timestamp; The initial speech-text segment is input into a large language model for semantic coherence analysis; Based on the results of semantic coherence analysis, semantically related continuous speech text segments are merged into semantically complete speech units, and the accurate start and end timestamps of the semantically complete speech units in the long video sample are determined. Based on the accurate start and end timestamps, the long video sample is split into corresponding short video segments.
4. The method for automatically generating short videos based on multi-level tag matching according to claim 3, characterized in that, Based on the results of semantic coherence analysis, semantically related continuous speech-text segments are merged into semantically complete speech units, and the accurate start and end timestamps of the semantically complete speech units in the long video sample are determined, including: Obtain multiple consecutive semantically related speech-text segments output by a large language model; Obtain the time interval between any two adjacent speech-text segments in a series of semantically related speech-text segments; When the time interval between any two adjacent speech text segments is less than the preset duration, multiple consecutive semantically related speech text segments will be merged into one speech unit. The audio unit is treated as a short video clip.
5. The method for automatically generating short videos based on multi-level tag matching according to claim 2, characterized in that, The method involves analyzing visual feature changes between consecutive video frames to detect scene transition points, thereby obtaining multiple short video segments after segmentation, including: Extract the continuous video frame sequence of the long video sample; Calculate the visual feature difference between any two consecutive adjacent video frames; When the visual feature difference between two consecutive adjacent video frames is greater than or equal to the preset visual feature difference threshold, the timestamp of the previous video frame is used as the scene transition point. Based on all detected scene transition points, the long video sample is divided into multiple scene segments; Each scene segment is treated as a short video clip.
6. The method for automatically generating short videos based on multi-level tag matching according to claim 1, characterized in that, The process of acquiring multiple long video samples, splitting each long video sample into multiple short video segments, and attaching a semantic tag set consisting of multiple semantic tags to each short video segment further includes: Select a short video clip and create a set of semantic tags corresponding to the short video clip; The main content in the short video clip is identified by a visual analysis model to generate at least one semantic tag, and the at least one semantic tag is included in the semantic tag set corresponding to the short video clip. The text content corresponding to the short video segment is analyzed by a text analysis model to generate at least one semantic tag, and the at least one semantic tag is included in the semantic tag set corresponding to the short video segment. The generated semantic tags form a hierarchical or parallel relationship with other existing semantic tags in the semantic tag set.
7. The method for automatically generating short videos based on multi-level tag matching according to claim 1, characterized in that, The process involves calculating the matching degree between each script segment and each short video segment based on the semantic tag set of each script segment and the semantic tag set of each short video segment in the video segment library. Then, based on the matching degree, one or more target short video segments are matched from the video segment library for each script segment, including: Retrieve a script fragment and a short video fragment to be matched; The semantic tag set of the script fragment and the semantic tag set to be matched are respectively converted into corresponding feature vectors; Calculate the duration matching factor by combining the estimated duration of the script segment with the actual duration of the short video segment; Based on the feature vector and the duration matching factor, a weighted calculation is performed to obtain the comprehensive matching score of the short video segment.
8. The method for automatically generating short videos based on multi-level tag matching according to claim 7, characterized in that, The method of calculating the matching degree between each script segment and each short video segment based on the semantic tag set of each script segment and the semantic tag set of each short video segment in the video segment library, and matching one or more target short video segments from the video segment library for each script segment according to the matching degree, further includes: Select a script segment and obtain the target number and score threshold of the corresponding short video segment; Determine whether the number of segments to be matched first is greater than or equal to the number of target short video segments; If the number of segments to be matched first is greater than or equal to the number of target short video segments, then the N short video segments with the highest overall matching score are selected as the target short video segments to match the script segment; N is a positive integer and N is equal to the number of target short video segments; If the number of segments to be matched first is less than the number of short video segments to be matched, then supplementary matching is performed from several short video segments that are the same as the script segment in terms of semantic tags at a lower level. The segments to be matched first and the short video segments to be matched together are used as the target short video segments to match the script segment.
9. The method for automatically generating short videos based on multi-level tag matching according to claim 8, characterized in that, The priority matching segment is the short video segment that is completely identical to the script segment in terms of higher-level semantic tags.
10. The method for automatically generating short videos based on multi-level tag matching according to claim 1, characterized in that, The process of aggregating all matched target short video segments according to the semantic coherence order of the script segments to generate a complete short video includes: Smoothness optimization; The smoothness optimization process includes at least one or more of the following: adding transition effects between adjacent target short video segments, aligning the scene transition points of the target short video segments with the beat points of the background music, and applying fade-in and fade-out effects to the audio track of the target short video segments.