Automated intelligent editing method for clothing live broadcast

CN122802735APending Publication Date: 2026-09-22卓尚服饰(杭州)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610719864.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]1.人工剪辑效率极低,规模化生产能力缺失:传统剪辑依赖人工逐帧观看、逐句聆听筛选商品讲解内容,手动完成切片、字幕制作、渲染与拼接,单条视频耗时长达数小时,人力成本高、出错率高,无法满足商家批量产出短视频的业务需求,难以适配直播电商的快节奏运营模式

Benefits of technology

[0040]本发明,通过智能Agent全流程调度与并行批处理架构,将传统人工单条视频数小时的剪辑时长压缩至分钟级,整体处理效率提升40%以上,结构化批量大模型推理使单批次处理吞吐量提升3倍,整体处理耗时缩短50%,断点续跑与中间结果复用机制使重复生成同一视频时处理耗时缩短70%以上,彻底解决了传统剪辑规模化生产能力缺失的痛点,可支撑每日数百条短视频的工业化批量产出,完美适配直播电商快节奏运营模式;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802735A_ABST
    Figure CN122802735A_ABST
Patent Text Reader

Abstract

The application discloses an automatic intelligent editing method for clothing live broadcast, and belongs to the technical field of intelligent video processing. The method performs audio segmentation processing on the original video of clothing live broadcast through intelligent Agent scheduling, generates sentence-level recognition results with timestamps through automatic speech recognition, and maps the results to the absolute time axis of the video. The method adopts programmed compliance filtering and Agent intelligent agent filtering mechanism based on large language models to eliminate prohibited, sensitive, promotional, interactive and non-commodity explanation content, and accurately retains clothing commodity explanation text. The method extracts highlight selling points based on large language models to form a highlight word set, and completes accurate video frame-level slicing according to the absolute time axis. The application can completely solve the problems of low efficiency, high compliance risk, subtitle word breaking, abnormal slicing picture and weak expression of selling points in traditional clothing live broadcast editing, and realizes end-to-end automation, industrialization and high-quality batch generation of clothing live broadcast short videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent video processing technology, specifically relating to an automated intelligent editing method for live streaming of clothing. Background Technology

[0002] Live-streaming e-commerce has become a core channel for apparel sales. Apparel live-streaming videos are characterized by their length, strong conversational style, dense interactive dialogue, a mix of promotional information and product explanations, and loose text structure. Editing hours-long apparel live-streaming videos into short product explanation videos that meet the requirements of short-video platforms is a core need for apparel e-commerce content operations. However, current traditional apparel live-streaming video editing techniques have many insurmountable shortcomings, specifically as follows:

[0003] 1. Manual editing is extremely inefficient and lacks large-scale production capacity: Traditional editing relies on manual viewing frame by frame and listening to sentence by sentence to select product explanation content, and manually completing slicing, subtitle production, rendering and splicing. A single video takes up to several hours, with high labor costs and high error rates. It cannot meet the business needs of merchants to produce short videos in batches and is difficult to adapt to the fast-paced operation mode of live e-commerce.

[0004] 2. Insufficient content compliance filtering capabilities and extremely high risk of platform penalties: Live streams are full of prohibited words, sensitive words, promotional language, and interactive codes to guide orders. Traditional technologies rely solely on keyword matching or manual screening, which cannot completely eliminate prohibited content. This leads to short videos being restricted or removed by the platform after publication, and may even result in account penalties, seriously affecting merchants' operations.

[0005] 3. Unreasonable subtitle segmentation logic, resulting in a poor user viewing experience: Traditional subtitles are segmented only according to a fixed number of characters or a fixed time, which easily leads to the core selling points of clothing being split in the middle, resulting in abrupt sentence breaks and incomplete information expression; at the same time, the subtitle display duration does not match the density of audio information, resulting in subtitles flashing too fast or lingering on the screen too slowly, which greatly reduces the user's viewing comfort and information reception efficiency.

[0006] 4. Insufficient video slicing precision and unstable final image quality: Traditional slicing uses a keyframe copying mode, which is prone to cutting in and out at non-keyframe positions, causing problems such as black screen, screen tearing, screen stuttering, and misalignment of subtitles and audio time, which seriously damages the visual integrity and professionalism of short videos.

[0007] 5. Lack of a mechanism for extracting unique selling points for apparel products, resulting in weak product conversion rates: Existing technologies lack the ability to extract selling points for apparel products. They cannot accurately identify core value information such as fabric material, pattern design, craftsmanship details, wearing scenarios, functional characteristics, and wearing experience. Subtitles and videos cannot enhance product selling points, making it difficult to attract user attention and improve product click-through and conversion rates.

[0008] 6. Lack of process reuse capability and high resource consumption costs: Traditional editing lacks the mechanism for resuming interrupted runs and reusing intermediate results. When repeatedly generating videos, the entire process must be executed from the beginning, which greatly wastes computing resources such as speech recognition and large model inference, increasing the overall cost of short video production.

[0009] Therefore, automated intelligent editing methods for live streaming of clothing are needed to solve the above problems. Summary of the Invention

[0010] The purpose of this invention is to provide an automated intelligent editing method for live streaming of clothing, in order to solve the problems mentioned in the background art.

[0011] To achieve the above objectives, the present invention provides the following technical solution: an automated intelligent editing method for live streaming of clothing, comprising the following steps:

[0012] The original video of the clothing live stream is obtained, and the audio stream of the original video is segmented according to a preset duration to obtain multiple audio segments.

[0013] The intelligent agent is scheduled to perform automatic speech recognition on the audio segment to obtain sentence-level recognition results containing text content, start timestamp, and end timestamp, and the timestamp is mapped to the absolute timeline of the original video;

[0014] The sentence-level recognition results are subjected to two-layer prohibited content filtering. The first layer is a procedural compliance filtering based on dictionary and pattern matching, and the second layer is semantic discrimination based on a large language model of the clothing product description agent. After filtering, the set of retained texts of clothing product descriptions is obtained.

[0015] The retained text set is subjected to Agent-based intelligent extraction of highlighted selling points based on a large language model. After splitting the retained text into clauses according to punctuation, clothing selling point words in the original text are extracted and a set of highlighted words is formed by prioritizing long words and retaining non-overlapping words.

[0016] Based on the absolute timestamps of the reserved text set, the original video is sliced ​​at frame level precision to obtain multiple video segments;

[0017] Perform multi-boundary constraint splitting on the subtitle text corresponding to each video segment to ensure that the highlighted words remain intact and are not split.

[0018] Perform visual rendering on the highlighted words in the subtitles and complete the subtitle burning process;

[0019] Multiple video clips are grouped and spliced ​​according to the target duration of the short video to generate the final short video.

[0020] This solution constructs an end-to-end automated editing workflow, completing the entire process from long videos to short videos without human intervention, compressing the traditional editing time of several hours to minutes, and significantly improving editing efficiency; it achieves precise binding of voice and text with video images through absolute timeline mapping, avoiding the problem of misalignment between subtitles and images from the root; and the intelligent agent scheduling throughout the process ensures the continuity of the process, adapting to the industrialized mass production needs of clothing live streaming short videos.

[0021] As a preferred implementation, the programmatic compliance filtering includes static dictionary matching and dynamic interaction pattern preliminary matching. The static dictionary matching performs text inclusion matching and removes matching text based on a violation word library, a sensitive word library, a promotional word library, an interactive script library, and a casual conversation word library. The dynamic interaction pattern preliminary matching identifies and removes simple interactive guidance scripts such as "debit + number" and "debit + Chinese number" using regular expressions. The accurate identification and removal of complex interactive guidance scripts and non-product explanation content is completed by semantic discrimination of the clothing product explanation agent intelligent agent based on a large language model, which is the core protection link of text screening.

[0022] By setting up programmatic compliance filtering and employing a dual-dimensional "dictionary + regular expression" approach for rapid screening, explicit and simple invalid content can be removed in milliseconds without needing to call a large language model, significantly reducing computational costs. It also specifically identifies common simple interactive prompts in live streaming, initially covering some invalid content that is not related to product explanations. This reduces the burden on the subsequent semantic judgment of the clothing product explanation agent based on the large language model. As the core step in text filtering, the semantic judgment of the clothing product explanation agent based on the large language model can accurately identify complex interactive phrases, subtle promotional content, and casual chat information that cannot be covered by regular expressions.

[0023] As a preferred implementation, the semantic discrimination of the clothing product explanation agent based on the large language model adopts a structured batch processing mechanism, which encapsulates candidate text into structured data containing timestamp and text fields and inputs them into the model in batches; the model output retains the original text content and timestamp, and a backup retention strategy is enabled when the model call is abnormal, so as not to discard candidate text.

[0024] By setting up structured batch processing and anomaly protection strategies, structured batch input improves the inference efficiency of large models and reduces the time consumption of a single call; retaining the original text and timestamps ensures that the results can be traced back to the video timeline without loss, avoiding information loss; the anomaly protection strategy prevents the silent discarding of text caused by model failure, improves the stability and fault tolerance of the system engineering, and ensures that the editing process is not interrupted.

[0025] As a preferred implementation, the Agent-based intelligent extraction of selling points based on a large language model only extracts words, idioms, or phrases that exist in the original text, without rewriting or generating the text; the extracted clothing selling point vocabulary categories include fabric material, pattern design, craftsmanship details, wearing scenarios, functional characteristics, and wearing experience.

[0026] By setting up a pure extraction-based highlighting mechanism, we ensure that the highlighted words are completely consistent with the original text spoken by the anchor, without any information tampering or distortion; the six major categories fully cover the core value dimensions of clothing products, accurately locate the selling points that users care about, provide accurate data support for the highlighting of subtitles, and strengthen the expression of product value.

[0027] As a preferred implementation, the subtitle multi-boundary constraint splitting includes punctuation boundary segmentation, word segmentation boundary secondary segmentation, and highlighted word boundary correction in sequence; when the candidate segmentation point is located inside the highlighted word, the segmentation point is automatically drifted to the front or back boundary of the highlighted word to ensure the integrity of the highlighted word.

[0028] By setting up a triple boundary constraint splitting mechanism, the problem of clothing selling points being broken down is prevented from the source, ensuring that the subtitles are naturally segmented and semantically complete; the word segmentation boundary is adapted to colloquial expressions, and the highlighted word boundary correction takes into account the balance of subtitle length and the integrity of selling points, greatly improving the fluency and professionalism of subtitle reading.

[0029] As a preferred implementation, after the subtitles are split, the total display time of the subtitles is allocated based on the proportion of the number of characters in each subtitle segment to the total number of characters, so that the display rhythm of the subtitles is consistent with the density of the voice information.

[0030] By setting a character proportion time allocation mechanism, the subtitle display rhythm is completely synchronized with the anchor's speaking speed and information density, without any issues of being too fast or too slow; the linear allocation rule ensures that the subtitles do not overlap or have time gaps, and the duration allocation is precise and controllable, further optimizing the user viewing experience.

[0031] As a preferred implementation, differentiated visual rendering is performed on the highlighted words in the subtitles, and dual-link subtitle output mode is supported; in the differentiated visual rendering, short highlighted words with a length not exceeding a preset number of characters are rendered using a highlight color combined with an enlarged font size, and long highlighted words with a length exceeding a preset number of characters are rendered using a highlight color combined with an underline or outline; the dual-link output mode includes an ASS subtitle script rendering mode and a transparent PNG layer overlay rendering mode.

[0032] By setting up differentiated rendering and dual-link output, length-based differentiated rendering balances visual enhancement and layout stability. Short words are enlarged to enhance visual impact, while long words are underlined / outlined to avoid layout chaos. Dual-link rendering is compatible with hardware and software environments of different video processing engines, eliminating process failures caused by missing rendering components, and improving the stability and portability of the system for cross-platform deployment.

[0033] As a preferred implementation, the two-layer prohibited content filtering adopts a hierarchical progressive filtering and parallel batch processing architecture; first, a coarse screening is completed through procedural filtering, and then a fine screening is completed through semantic discrimination of the clothing product explanation agent intelligent agent based on a large language model; candidate texts are processed in parallel in batches, and when the model call fails, it automatically switches to the procedural filtering result as a backup.

[0034] By setting up a hierarchical and parallel batch processing architecture, a combination of "rapid coarse screening + precise fine screening" is achieved, resulting in a dual improvement in filtering efficiency and accuracy. Parallel batch processing significantly increases the inference throughput of large models and shortens the overall processing time. The dual-safety strategy further enhances system stability and ensures the normal operation of the editing process under extreme scenarios.

[0035] In a preferred embodiment, the frame-level precision slicing adopts a post-decoding re-encoding slicing method; the video segments are automatically grouped and spliced ​​according to a preset target duration interval, and if the cumulative duration does not reach the lower limit, the segments are forcibly merged; if the upper limit is reached, the segments are grouped and output to ensure that the final video length is compliant and the semantics are coherent.

[0036] By setting up decoding and recoding segments and automatic grouping and splicing, frame-level decoding and recoding segments completely eliminate problems such as black screen, screen tearing, and video stuttering, resulting in stable video quality. The preset target duration range adapts to the publishing requirements of all short video platforms, and automatic grouping and splicing preserves the semantic coherence of product descriptions without the need for manual adjustment of segment order, further improving the level of automation.

[0037] As a preferred implementation, the system also includes a breakpoint resume and intermediate result reuse step, and a scenario generalization and expansion step. The breakpoint resume and intermediate result reuse step involves the intelligent agent recording the processing status and intermediate results of each stage in real time and generating file hash values. When the process is restarted after an interruption, the validity of the intermediate results is verified by the hash value, and the completed processing stages are automatically skipped. The scenario generalization and expansion step involves updating the rule lexicon, adjusting the prompt words of the large language model, and modifying the selling point category configuration to quickly adapt to the live video editing needs of different product categories.

[0038] By setting up breakpoint resume and scenario generalization mechanisms, breakpoint resume can effectively avoid repeatedly executing time-consuming steps such as speech recognition and large model inference. When repeatedly generating the same video, the processing time is greatly reduced, and the cost of calling large models and automatic speech recognition is significantly reduced, thereby greatly reducing the overall cost of industrial production. Scenario generalization does not require the reconstruction of the core processing flow. It can achieve multi-scenario adaptation simply by configuring parameters, making it highly versatile and maximizing the value of commercial applications.

[0039] Compared with the prior art, the beneficial effects of the present invention are:

[0040] This invention, through intelligent agent full-process scheduling and parallel batch processing architecture, compresses the editing time of traditional manual single video, which used to take several hours, to minutes, improving overall processing efficiency by more than 40%. Structured batch large model inference increases the throughput of single batch processing by 3 times and reduces the overall processing time by 50%. The breakpoint resume and intermediate result reuse mechanism reduces the processing time when repeatedly generating the same video by more than 70%. It completely solves the pain point of the lack of large-scale production capacity in traditional editing, and can support the industrial batch production of hundreds of short videos per day, perfectly adapting to the fast-paced operation mode of live e-commerce.

[0041] This invention employs a two-layer prohibited content filtering mechanism: "programmatic coarse screening + large language model agent intelligent screening." This reduces platform penalty risks by over 99%. Static dictionary matching and preliminary pattern matching using regular expressions can remove explicit and simple invalid content in milliseconds. The semantic discrimination of the clothing product explanation agent intelligent body based on the large language model serves as the core protection link in text filtering. It can accurately identify complex interactive guidance scripts, subtle promotional scripts, and casual conversation content. The overall non-product explanation content recall rate is no less than 99%, and the semantic discrimination accuracy of the large language model agent intelligent body is no less than 95%. It can accurately identify and remove illegal, sensitive, promotional, interactive, and casual conversation content. The exception protection strategy prevents the loss of effective text due to model failure, increasing the system's fault tolerance rate to 100% and ensuring the compliance of short video content from the source.

[0042] This invention improves user information reception efficiency by more than 80% through a triple boundary constraint subtitle splitting and character proportion time allocation mechanism, and the accuracy of subtitle segmentation is no less than 99%. It fundamentally avoids the problem of core selling points of clothing being broken down. The linear duration allocation based on character proportion makes the subtitle display rhythm completely synchronized with the anchor's speaking speed and voice information density, completely solving the problem of traditional subtitles flashing too fast or lingering on the screen too slow. Differentiated highlight rendering increases the user attention of short videos by more than 30%, greatly improving user viewing comfort and information reception efficiency.

[0043] This invention features a pure extraction mechanism for identifying unique selling points in apparel. It boasts a highlight word extraction accuracy rate of at least 98%, accurately identifying six core selling points, including fabric material, pattern design, and craftsmanship details. A differentiated visual rendering strategy dynamically matches the rendering style based on the length of the highlighted words. Short words are enlarged to enhance visual impact, while long words are underlined / outlined to maintain layout stability. This enhanced visualization of selling points effectively highlights the core value of the product, improving the commercial appeal of short videos and significantly increasing click-through rates and conversion rates. It also solves the problem of weak selling point expression in traditional editing.

[0044] This invention employs a post-decoding re-encoding frame-level precision slicing technology, achieving 100% stability in video clips. This completely eliminates problems such as black screens, screen tearing, and stuttering caused by traditional keyframe copying and slicing, reaching professional broadcast-grade standards. The timeline mapping accuracy reaches millisecond level, fundamentally solving the industry problem of misalignment between subtitles, audio, and video. The automatic grouping and splicing algorithm ensures that the final video length is strictly controlled within the platform standard range of 15-30 seconds, with semantic coherence, eliminating the need for manual secondary adjustments, and ensuring stable and controllable final video quality.

[0045] This invention features a dual-link subtitle output mode compatible with various deployment environments such as Windows, Linux, and cloud containers, achieving 100% cross-platform stability. It eliminates process failures caused by missing rendering components. Its hierarchical progressive filtering and parallel batch processing architecture allows for editing even in extreme scenarios where large language models are completely unavailable, through programmatic filtering results. The scene generalization and extension mechanism eliminates the need to reconstruct the core processing flow, enabling rapid adaptation to live streaming editing needs across all categories, including beauty, home furnishings, 3C products, and food, simply through parameter configuration, making it highly versatile. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the overall process of the present invention. Detailed Implementation

[0047] The present invention will be further described below with reference to embodiments.

[0048] The following embodiments are used to illustrate the present invention, but should not be used to limit the scope of protection of the present invention. The conditions in the embodiments can be further adjusted according to specific conditions, and simple improvements to the method of the present invention under the premise of the concept of the present invention are all within the scope of protection claimed by the present invention.

[0049] The present invention will be further described below with reference to embodiments. These embodiments are for illustrative purposes only and should not be construed as limiting the scope of protection of the present invention. The conditions in the embodiments may be further adjusted according to specific conditions, and simple modifications to the method of the present invention under the premise of the present invention's concept are all within the scope of protection claimed by the present invention.

[0050] Please see Figure 1 This invention provides an automated intelligent editing method for live streaming of clothing, the specific implementation of which is as follows:

[0051] Example 1

[0052] This embodiment provides the basic workflow of an automated intelligent editing method for live streaming of clothing. This method is uniformly scheduled and executed by an intelligent agent without human intervention, and specifically includes the following steps:

[0053] S1. Audio Segmentation Processing: This step is achieved collaboratively by the audio extraction unit, the segmentation execution unit, and the status recording unit. The system first separates the pure audio stream from the original video through the audio extraction unit. The segmentation execution unit evenly divides the audio stream according to a preset duration L (preferably 10 minutes) to generate continuous audio segments, while recording the sequence number k, start time offset, and end time offset of each segment. The status recording unit marks the processing status of each segment in real time. When the process is interrupted and restarted, it automatically resumes execution from the segment that has been processed. This step realizes the lightweight splitting of long audio and the function of resuming interrupted processing, effectively reducing the processing pressure of a single speech recognition, avoiding the interruption of the entire process caused by the failure of long audio recognition, and improving the overall processing efficiency by more than 40%.

[0054] S2. Sentence-level speech recognition and timeline mapping: This step is achieved collaboratively by the intelligent agent scheduling unit, the automatic speech recognition interface, and the timeline mapping calculation unit. The intelligent agent scheduling unit sequentially calls the automatic speech recognition interface according to the generation order of the audio segments, performs speech recognition processing on each audio segment, and generates a sentence-level recognition result containing the text content, the segment's start timestamp τ_start_i, and the segment's end timestamp τ_end_i. The timeline mapping calculation unit superimposes the segment timestamp of each sentence-level recognition result with the time offset of the corresponding audio segment, converts it into the absolute timestamp of the original video, and establishes a precise time correspondence between the speech text and the video frame.

[0055] Absolute time axis mapping formula:

[0056]

[0057]

[0058] in:

[0059] The absolute start time of the i-th recognized sentence in the original video is directly calculated using the formula.

[0060] The absolute end time of the i-th recognized sentence in the original video is directly calculated using the formula.

[0061] The current audio segment number, starting from 1 and incrementing sequentially;

[0062] The preset duration of a single audio segment is a system configuration parameter, with a preferred value of 10 minutes.

[0063] The start time of the i-th recognized sentence within the current audio segment is returned by the automatic speech recognition interface;

[0064] : The end time of the i-th recognized sentence in the current audio segment, returned by the automatic speech recognition interface.

[0065] This step enables sentence-level text generation with absolute timestamps, providing a unique time positioning basis for subsequent video slicing and subtitle synchronization. The time alignment accuracy reaches the millisecond level, fundamentally solving the industry problem of misalignment between subtitles, audio, and video.

[0066] S3. Two-layer prohibited content filtering: This step is achieved collaboratively by a programmatic filtering unit, a clothing product description agent based on a large language model, and a result integration unit. The programmatic filtering unit first performs a rapid coarse screening on the sentence-level recognition results, eliminating explicit, simple, and invalid content. The clothing product description agent based on a large language model performs precise semantic discrimination on the remaining candidate text, retaining only content highly relevant to the clothing product description; this is the core protection step in text screening. The result integration unit merges the valid texts to form a set of retained clothing product description texts. This step achieves precise filtering of pure product description content, comprehensively eliminating invalid information such as violations, sensitive content, promotions, interactions, and casual chat, ensuring the compliance of short video content from the source, and reducing the risk of platform penalties by more than 99%.

[0067] Programmatic compliance filtering matching function:

[0068]

[0069] in:

[0070] : The dictionary matching flag for the i-th recognized sentence, where 1 indicates an invalid content match and 0 indicates no match;

[0071] : The text content of the i-th recognized sentence;

[0072] The excluded word set is a combination of the violation word library, the sensitive word library, the promotion word library, the interactive language library, and the casual conversation word library.

[0073] Semantic discrimination function of clothing product explanation agent based on large language model:

[0074]

[0075]

[0076] in:

[0077] The set of candidate sentences after procedural filtering;

[0078] The set of text retained after semantic discrimination;

[0079] The result of the clothing product description agent based on the large language model judging the candidate sentence c, 1 indicates retention, 0 indicates rejection.

[0080] S4. Intelligent Extraction of Highlighted Selling Points Based on a Large Language Model: This step is achieved collaboratively by a clause segmentation unit, an extractive selling point recognition unit, and a highlight word deduplication unit. The clause segmentation unit divides the retained text into multiple independent clause units according to punctuation marks; the extractive selling point recognition unit extracts clothing-specific selling point vocabulary from the clause units; the highlight word deduplication unit processes the extracted vocabulary according to the rules of prioritizing longer words and retaining non-overlapping words, ultimately forming a set of highlighted words.

[0081] Extracted selling point constraints:

[0082]

[0083] in:

[0084] : The collection of clauses after the text is split according to punctuation;

[0085] A collection of key selling words extracted from the short sentence u;

[0086] The collection of all words, idioms, or phrases contained in clause u.

[0087] This constraint ensures that the highlighted words are derived solely from the original text, without any rewriting or generation, guaranteeing 100% information authenticity. This step enables the precise extraction of the core selling points of the clothing, providing a data foundation for subsequent subtitle highlighting and rendering, effectively enhancing the expression of product value and increasing the commercial appeal of the short video.

[0088] S5. Frame-Level Precise Video Slicing: This stage is achieved collaboratively by the frame-level decoding unit, the re-encoding slicing unit, and the timestamp verification unit. The frame-level decoding unit decodes the original video completely based on the absolute timestamps corresponding to the retained text set. The re-encoding slicing unit segments the decoded video stream at precise frame positions. The timestamp verification unit calibrates the slice time in real time to eliminate time drift errors. This stage enables the generation of video clips without any abnormal visuals, completely eliminating issues such as black screens, screen tearing, and stuttering, achieving 100% stability in the video clips.

[0089] S6. Subtitle Multi-Boundary Constraint Segmentation: This step is achieved collaboratively by the punctuation segmentation unit, the word segmentation unit, and the highlighted word boundary correction unit. The three units sequentially perform three-level segmentation processing on the subtitle text to ensure that the segmentation points do not fall inside the highlighted words. This step realizes a scientific and reasonable subtitle segmentation function, ensuring that the subtitles are naturally broken and that the selling points are intact and not split, which greatly improves the reading fluency of the subtitles.

[0090] S7. Highlighted Word Rendering and Subtitle Burning: This step is achieved collaboratively by the rendering strategy matching unit and the dual-link burning unit. The rendering strategy matching unit matches the corresponding rendering style according to the length of the highlighted words. The dual-link burning unit adds the rendered subtitles to the video clip. This step realizes the visualization enhancement function of selling point words, effectively highlighting the core selling points of the product and improving the visual appeal of the short video.

[0091] S8. Automatic Grouping and Stitching of Video Clips: This step is achieved collaboratively by the duration statistics unit, the grouping calculation unit, and the temporal stitching unit. The duration statistics unit counts the duration of each video clip; the grouping calculation unit groups the video clips according to the target duration constraints of the short video platform; the temporal stitching unit stitches the video clips in the same group into a complete short video clip according to the original temporal sequence. This step realizes the automatic generation function of short video clips that meet the platform requirements. The clip duration is fully compliant and semantically coherent, without the need for manual secondary adjustments.

[0092] The method in this embodiment, through the full-process scheduling of intelligent agents, realizes the automated editing of long videos from clothing live streams into short videos, compressing the traditional editing time of several hours to minutes, and significantly improving editing efficiency. At the same time, through mechanisms such as timeline mapping, double-layer filtering, and subtitle optimization, it effectively solves the problems of high compliance risks, easy word breaks in subtitles, and unstable images that exist in traditional editing. The generated short videos have high compliance, highlight selling points, smooth subtitles, and stable images, which are suitable for the industrialized mass production needs of clothing live stream short videos.

[0093] Example 2

[0094] The difference between this embodiment and Embodiment 1 is that the procedural compliance filtering step of the two-layer prohibited content filtering in step S3 has been refined, as follows:

[0095] The programmatic compliance filtering process is achieved collaboratively by a static dictionary, a regular expression matching engine, and a text removal execution unit. The static dictionary is pre-built and loaded with five categories of dictionaries: a violation dictionary, a sensitive dictionary, a promotional dictionary, an interactive script dictionary, and a casual conversation dictionary. The text removal execution unit performs string inclusion matching between sentence-level identified text and words in the dictionaries. If the text matches any word in any dictionary, it is directly determined as invalid content and removed. The regular expression matching engine loads preset regular expression rules and performs preliminary pattern matching on the text. It identifies and removes simple interactive guidance scripts in the form of "debit + number" or "debit + Chinese number" commonly seen in live streaming scenarios. For complex interactive guidance scripts, obscure promotional scripts, and casual conversation content that cannot be covered by regular expressions, a clothing product explanation agent based on a large language model performs accurate semantic recognition and removal. This is the core protection step of text screening and can cover more than 99% of invalid content that is not related to product explanations.

[0096] Interactive pattern regular expression:

[0097]

[0098] The static dictionary supports dynamic updates. When the platform releases the latest violation rules, words can be added or deleted directly in the dictionary without modifying the core processing logic. The regular expression rules can cover all mainstream interactive guidance scripts, with a matching recall rate of no less than 99%.

[0099] This step enables rapid filtering of explicit, simple, and invalid content, completing the matching and removal of individual texts in milliseconds without needing to call a large language model, significantly reducing computational costs. Simultaneously, it specifically identifies simple interactive guidance scripts unique to live streaming scenarios, initially covering some invalid content that is not related to product descriptions. This alleviates the processing pressure on the subsequent semantic judgment of the clothing product description agent based on the large language model. The clothing product description agent based on the large language model, as the core filtering step, solves the industry pain point that regular expressions cannot identify complex and obscure non-product description content.

[0100] Example 3

[0101] The difference between this embodiment and Embodiment 1 is that the semantic discrimination step of the clothing product explanation agent based on a large language model for the two-layer prohibited content filtering in step S3 has been refined, as follows:

[0102] The semantic discrimination stage of the clothing product explanation agent based on the large language model is achieved collaboratively by a structured data encapsulation unit, a batch inference unit, and an anomaly protection unit. The structured data encapsulation unit encapsulates the candidate texts, after programmatic filtering, into JSON arrays containing timestamp and text fields. The batch inference unit divides the encapsulated structured data into multiple batches according to a fixed batch size B (preferably 32 entries / batch), and inputs them in parallel into the clothing product explanation agent based on the large language model for inference. The clothing product explanation agent based on the large language model judges each candidate text according to preset clothing product explanation relevance judgment rules, retaining only text content highly relevant to the clothing product explanation, and outputting the original text content and corresponding timestamp information without any text rewriting or modification. The anomaly protection unit monitors the calling status of the clothing product explanation agent based on the large language model in real time. When a model call fails, returns a null value, or the return format does not meet requirements, it automatically retains all candidate texts in that batch without silent discarding.

[0103] Batch inference parallel processing formula:

[0104]

[0105]

[0106] in:

[0107] : The p-th candidate text batch;

[0108] The semantic discrimination result of the p-th batch;

[0109] Total number of batches;

[0110] The final set of retained text.

[0111] The pre-defined relevance judgment rules are based on the business scenario of clothing live streaming. They focus on identifying texts containing relevant content such as clothing fabric, pattern, craftsmanship, matching, function, and experience, while removing non-product explanation content such as casual chat, interaction, and promotion. The judgment accuracy rate is no less than 95%.

[0112] This step enables precise semantic filtering of clothing product description text. The efficiency of structured batch processing is more than 50% higher than that of single-item processing, effectively reducing the calling cost of large language models. The exception protection strategy eliminates the loss of effective text caused by model failure, increasing the system's fault tolerance rate to 100%, ensuring the continuity of the editing process, and avoiding the interruption of the entire process due to model problems.

[0113] Example 4

[0114] The difference between this embodiment and Embodiment 1 is that the intelligent extraction of agent-highlighted selling points based on a large language model in step S4 has been refined, as follows:

[0115] The Agent-based intelligent extraction of highlighted selling points, based on a large language model, is achieved collaboratively by a clause splitting unit, a selling point classification and recognition unit, and a long-word-first deduplication unit. The clause splitting unit divides the retained text set into multiple independent clause units according to punctuation marks such as commas, periods, exclamation marks, and question marks. The selling point classification and recognition unit uses an extraction algorithm to extract words, idioms, or phrases from the original text within the clause units without any text rewriting, generation, or expansion. The extracted selling point words are divided into six categories: fabric material, pattern design, craftsmanship details, wearing scenarios, functional characteristics, and wearing experience. The long-word-first deduplication unit arranges the extracted selling point words in descending order of length and matches them sequentially with the text. If the matching interval of a word overlaps with the interval of the retained highlighted words, only the longer word is retained, ultimately generating a non-overlapping, fully covered set of highlighted words.

[0116] Long word priority deduplication rule:

[0117]

[0118] in:

[0119] : The character length of the i-th highlighted word, sorted in descending order.

[0120] The six major selling point categories can be dynamically adjusted according to the clothing category. For example, for children's clothing live streaming, selling point categories related to children, such as "soft and anti-friction", can be added to improve the targeting of the extraction.

[0121] This step enabled the pure text extraction of the core selling points of clothing, ensuring that the highlighted words were completely consistent with the original text spoken by the anchor, without any information tampering or distortion; the six major categories fully covered the core value dimensions of clothing products, accurately locating the selling points that users care about; the deduplication rule of prioritizing long words avoided the problem of short words covering long words, ensuring the integrity of the selling points expression, and the accuracy rate of highlighted word extraction was no less than 98%, providing accurate data support for subsequent subtitle highlighting and rendering.

[0122] Example 5

[0123] The difference between this embodiment and Embodiment 1 is that the subtitle multi-boundary constraint decomposition step in step S6 has been refined, as follows:

[0124] The subtitle multi-boundary constraint splitting process is implemented collaboratively by a punctuation segmentation unit, a length threshold configuration unit, a word segmentation unit, and a highlighted word boundary correction unit. The punctuation segmentation unit first removes punctuation marks and invalid whitespace characters from the end of the subtitle text. Then, it performs a first-level segmentation based on punctuation marks such as commas, periods, exclamation marks, and question marks, generating preliminary subtitle segments. The length threshold configuration unit presets a maximum character count threshold θ_len for subtitle segments (preferably 16 characters). For subtitle segments with character lengths exceeding the preset threshold, the word segmentation unit performs a second segmentation based on a Chinese word segmentation model, selecting word boundaries near the midpoint of the text to ensure that each subtitle segment has a suitable length. The highlighted word boundary correction unit detects the position of all candidate segmentation points. If a segmentation point falls inside a highlighted word, it automatically drifts to the front or back boundary of the highlighted word, selecting the position that minimizes the length difference between the left and right subtitle segments as the final segmentation point.

[0125] Word segmentation boundary secondary segmentation objective function:

[0126]

[0127]

[0128] in:

[0129] Optimal word segmentation position;

[0130] : The cumulative character length of the first k words;

[0131] : The total number of characters in the j-th sub-segment after the first level of segmentation;

[0132] The p-th word output by the Chinese word segmentation model;

[0133] Total word count of this sub-section.

[0134] Highlighted word boundary correction function:

[0135]

[0136] Preferred parameters:

[0137] , ;

[0138] Constraints:

[0139] The segmentation point must be moved to the front or back boundary of the highlighted word; segmentation must not be performed inside the highlighted word.

[0140] In this embodiment, the preset maximum character count threshold is preferably 16 characters, which can be adaptively adjusted according to the width of the subtitle display.

[0141] This step achieves a scientific and reasonable subtitle segmentation function. Punctuation boundary segmentation adapts to the sentence break logic of colloquial expressions, making the subtitles' sentence breaks natural; word segmentation boundary segmentation avoids the problem of reading long subtitle segments and ensures the balanced length of each subtitle segment; highlight word boundary correction fundamentally avoids the problem of clothing selling point words being broken up, ensuring the integrity of information expression; in this embodiment, the sentence breakation accuracy of subtitle segmentation is no less than 99%, effectively improving the reading fluency of subtitles and increasing user information reception efficiency by more than 60%.

[0142] Example 6

[0143] The difference between this embodiment and Embodiment 1 is that the subtitle time allocation step in step S6 has been refined, as follows:

[0144] The subtitle time allocation process is achieved collaboratively by a character statistics unit, a duration calculation unit, and a timing verification unit. The character statistics unit counts the number of characters in each subtitle segment and the total number of characters in all subtitle segments. The duration calculation unit accurately allocates the total display duration of the subtitles based on the proportion of the number of characters in each subtitle segment to the total number of characters. The timing verification unit verifies the allocated subtitle timing to ensure that there is no time overlap or time gap.

[0145] Subtitle display time allocation formula:

[0146]

[0147] in:

[0148] The start time of the display of the i-th subtitle segment is calculated directly by the formula;

[0149] The display end time of the i-th subtitle segment is calculated directly using the formula.

[0150] : The display duration of the i-th subtitle segment;

[0151] The total display duration of the original subtitle text is obtained directly from the difference between the end timestamp and the start timestamp of the corresponding sentence-level text obtained through automatic speech recognition.

[0152] The number of characters in the k-th subtitle segment is obtained by directly counting the subtitle segments after splitting them using a text character counting tool.

[0153] The total number of characters in all subtitle segments is calculated by summing the character counts of all subtitle segments. ,in This represents the total number of subtitle segments;

[0154] : The sequence number of the subtitle segment, generated sequentially starting from 1 according to the segmentation order, and is a positive integer.

[0155] Control logic:

[0156] This formula uses a purely linear allocation logic, automatically allocating the total display duration according to the proportion of characters in the subtitle segment to the total number of characters. No manual intervention or parameter adjustment is required; the system automatically completes all calculations and timing allocations.

[0157] Constraints:

[0158] Total duration conservation: The sum of the display durations of all subtitle segments is strictly equal to the total duration. ,Right now ;

[0159] No time overlap: There is no time overlap between any two adjacent subtitle segments, that is, for any... All satisfy ;

[0160] No time gap: There is no time gap between any two adjacent subtitle segments, that is, for any... All satisfy ;

[0161] Segmentation safety: All segmentation points strictly adhere to the triple constraints of punctuation boundaries, word segmentation boundaries, and highlighted word boundaries, and segmentation is not allowed within highlighted words.

[0162] Technical benefits: This allocation mechanism ensures that the display rhythm of the subtitles is completely synchronized with the speaker's speed and the density of audio information, thus completely solving the problems of traditional subtitles flashing too fast or lingering on the screen too slowly; users can easily and completely receive product information, improving viewing comfort and information reception efficiency by more than 80%; at the same time, the linear allocation rule ensures the accuracy of the subtitle timing and avoids the problem of subtitles and audio being out of sync.

[0163] This feature enables precise allocation of subtitle display duration, with millisecond-level accuracy, ensuring a perfect match between subtitle rhythm and audio information density, thus further optimizing the user's viewing experience.

[0164] Example 7

[0165] The difference between this embodiment and Embodiment 1 is that the highlighted word differentiation rendering and subtitle burning process in step S7 has been refined, as follows:

[0166] The differentiated rendering of highlighted words and the subtitle burning process are achieved collaboratively by the length judgment unit, the rendering style execution unit, the style isolation unit, the ASS script generation unit, the PNG transparent layer rendering unit, and the video processing engine. The length judgment unit identifies the character length of each highlighted word, and the rendering style execution unit adopts different rendering methods according to the length of the highlighted word.

[0167] Differentiated rendering style functions:

[0168] like

[0169]

[0170] in:

[0171] Highlight the word range;

[0172] : Character length of the highlighted word;

[0173] The preferred length threshold for short highlighted words is 4 characters.

[0174] The style stacking operator indicates that multiple rendering effects are applied simultaneously.

[0175] Highlight color rendering;

[0176] Font size increased by 1.2 times;

[0177] Underline rendering;

[0178] Outline rendering.

[0179] The style isolation unit achieves rendering isolation through placeholder injection and style tag replacement mechanism. First, the highlighted words are replaced with unique placeholders. After the overall subtitle layout is completed, the placeholders are replaced with highlighted words with style tags to avoid style tags interfering with text matching.

[0180] To adapt to different video processing environments, this section provides two subtitle burning links:

[0181] The first type is the ASS subtitle script rendering link, where the ASS script generation unit generates an ASS subtitle script containing Style definition and Dialogue event, and the subtitles are burned into the video through the subtitles filter of the video processing engine;

[0182] The second method is a transparent PNG layer overlay rendering chain. The PNG transparent layer rendering unit renders the subtitles as a PNG transparent layer with an alpha channel. Through the overlay filter and temporal compositing expression of the video processing engine, the subtitle layer is overlaid on the video screen.

[0183] Formula for creating transparent PNG layers:

[0184]

[0185]

[0186] in:

[0187] : The synthesized output video frames;

[0188] : Original video frames;

[0189] : The PNG transparent layer of the i-th subtitle segment;

[0190] : The time gating function for the i-th subtitle segment;

[0191] : The display time interval of the i-th subtitle segment.

[0192] The two link systems can automatically switch according to the deployment environment without manual configuration. They are compatible with multiple platforms such as Windows, Linux, and cloud containers, and cross-platform stability can reach 100%. The highlight color can be adjusted according to the overall color tone of the video. For example, for videos with light backgrounds, high-contrast colors such as red and yellow are used; for videos with dark backgrounds, prominent colors such as white and light colors are used to ensure the visibility of the highlight effect.

[0193] This step enables differentiated visualization enhancement of selling points. The length-adaptive rendering method balances visual impact and layout stability. Short words are enlarged to enhance visual appeal, while long words are underlined / outlined to avoid subtitle confusion. The dual-link burning mechanism ensures system compatibility in different deployment environments, eliminating process failures caused by missing rendering components. In this embodiment, the rendering accuracy of highlighted words is no less than 99%, effectively highlighting product selling points and increasing user attention to short videos by more than 30%.

[0194] Example 8

[0195] The difference between this embodiment and Embodiment 1 is that the hierarchical progressive filtering and parallel batch processing architecture for the two-layer prohibited content filtering in step S3 has been refined, as follows:

[0196] The two-layer prohibited content filtering adopts a hierarchical progressive filtering and parallel batch processing architecture, which is implemented collaboratively by a batch scheduling unit, a parallel execution unit, and a result switching unit. The batch scheduling unit first schedules the procedural filtering unit to complete the coarse screening, eliminating most of the explicitly invalid content; then, the remaining candidate text is divided into multiple batches according to a fixed batch size B (preferably 32 texts / batch), and the parallel execution unit sends the text from multiple batches in parallel to the clothing product explanation agent based on a large language model for inference; the result switching unit monitors the calling status of the clothing product explanation agent based on the large language model in real time, and when the model call fails, it automatically switches to the result of procedural filtering as the backup output.

[0197] This architecture implements a two-layer filtering mode of "rapid coarse screening + precise fine screening". The hierarchical processing method ensures both filtering efficiency and filtering accuracy. The parallel batch processing architecture increases the processing throughput by more than 3 times compared to single-processing, and reduces the overall processing time by 50%. The dual-safety strategy further enhances the stability of the system. Even in extreme scenarios where the large language model is completely unavailable, the editing process can still be completed through the results of programmatic filtering, ensuring business continuity.

[0198] Example 9

[0199] The difference between this embodiment and Embodiment 1 is that the frame-level video slicing in step S5 and the video segment grouping and splicing in step S8 have been refined, as follows:

[0200] The frame-level video slicing process is achieved collaboratively by a decoding unit, a re-encoding unit, and a timestamp calibration unit. The decoding unit decodes the original video completely based on the absolute timestamps corresponding to the retained text set, converting the video stream into frame-by-frame pixel data. The re-encoding unit segments the pixel data at precise frame positions and then re-encodes it to generate independent video segments. The timestamp calibration unit calibrates the timestamps of each video segment to ensure that the time error of the slice does not exceed one frame. Compared with traditional keyframe copy slicing, this method can effectively eliminate the picture anomalies caused by non-keyframe slicing, and the picture stability of the video segments can reach 100%, meeting professional broadcast-grade standards.

[0201] The video clip grouping and splicing process is achieved collaboratively by a duration statistics unit, a grouping calculation unit, and a timing splicing unit; the duration statistics unit calculates the precise duration of each video clip. The group calculation units are based on a target duration range of 15 to 30 seconds. (Preferred) Second, The video clips are grouped in seconds.

[0202] Automatic grouping and splicing algorithm:

[0203] Initialize the current group Cumulative duration ;

[0204] Iterate through each video segment in its original order. :

[0205] a. If Then add the fragment to the current group and update. ;

[0206] b. Otherwise, if If the current group is selected as a candidate final film, a new group will be created. ,renew ;

[0207] c. If Then force Join the current group, output the group, and start a new group. ,renew ;

[0208] After the traversal is complete, if the current group is not empty and If the condition is met, output the current group; otherwise, merge it into the previous group.

[0209] The temporal splicing unit stitches together video clips from the same group into a complete short video clip according to their original temporal sequence. The grouping and splicing process preserves the semantic coherence of the product description, without forcibly splitting up the continuous explanatory content. The length of the final video fully meets the platform's publishing requirements, requiring no manual adjustments.

[0210] This step enables the automatic generation of short videos with stable visuals and compliant duration, completely solving the problems of abnormal visuals and inconsistent durations in traditional editing, improving the quality and platform compatibility of the finished product, and maximizing industrial production efficiency.

[0211] Example 10

[0212] The difference between this embodiment and Embodiment 1 is that it adds a mechanism for resuming execution from breakpoints, reusing intermediate results, and generalizing and extending scenarios, as detailed below:

[0213] The mechanism for resuming execution from breakpoints and reusing intermediate results is implemented collaboratively by a state detection unit, a hash verification unit, and an intermediate result storage unit. The intermediate result storage unit stores intermediate results from each stage of the editing process in real time, including audio segmentation results, speech recognition results, filtered text sets, and highlighted word sets, while generating a file hash value for each intermediate result. When the process starts, the state detection unit first checks the intermediate results in the output directory. The hash verification unit verifies the consistency between the input video and the intermediate results using the file hash values. If the intermediate results are valid, the completed processing steps are automatically skipped, and the process proceeds directly to the next step. If the intermediate results are invalid or missing, the process continues from the breakpoint. This mechanism effectively avoids repeatedly executing time-consuming steps such as speech recognition and large model inference. When generating the same video repeatedly, the processing time is reduced by more than 70%, and the cost of calling large models and automatic speech recognition is reduced by 80%, significantly reducing the overall cost of industrial production. At the same time, storing intermediate results also facilitates subsequent backtracking and modification. For example, when it is necessary to adjust the rendering style of highlighted words, the generated video clips and subtitle text can be reused directly without re-executing the entire process.

[0214] The scenario generalization and expansion mechanism is implemented collaboratively by the rule-based vocabulary configuration unit, the prompt word adjustment unit, and the selling point category configuration unit. By updating the violation and sensitive word libraries in the rule-based vocabulary configuration unit, it can adapt to the violation rules of different platforms. By adjusting the relevance judgment prompt words in the large language model in the prompt word adjustment unit, it can adapt to the filtering of product description texts for different categories such as beauty, home furnishings, 3C products, and food. By modifying the selling point categories in the selling point category configuration unit, it can extract the core selling point vocabulary of different categories. By adjusting the style parameters of subtitle rendering, it can adapt to the video style requirements of different platforms. For example, when adapting to beauty live-stream editing, the selling point category can be adjusted to "product ingredients, efficacy, usage method, suitable skin type, makeup effect, and makeup duration," etc., and the relevance judgment rules of the clothing product description agent based on the large language model can also be adjusted to focus on identifying beauty-related explanation content. When adapting to home furnishing live-stream editing, the selling point category can be adjusted to "material, size, function, installation method, applicable scenarios, and durability," etc., while simultaneously updating the violation vocabulary and interaction mode rules to adapt to the conversational scenarios of home furnishing live streams. This mechanism does not require restructuring of the core processing flow; it can be adapted to multiple scenarios simply by configuring parameters, making it highly versatile and maximizing its commercial application value.

[0215] The working principle and usage process of this invention are as follows: First, the original live-stream video of the clothing product is uploaded. The system automatically extracts the audio and segments it according to a preset duration. Then, an intelligent agent is scheduled to perform speech recognition, generating sentence-level text with timestamps and mapping it to an absolute timeline. Next, a two-layer prohibited content filtering process is executed. First, a simple coarse screening of invalid content is completed through programmatic filtering. Then, a core precise screening is completed through semantic discrimination by a clothing product explanation agent based on a large language model, eliminating invalid content and retaining the product explanation text. Next, intelligent extraction of highlighted selling points by the agent based on a large language model is performed, and frame-level video slicing is completed based on timestamps. Subtitles undergo triple boundary splitting and character proportion time allocation, and highlighted words are rendered differently and subtitles are burned through a dual-link process. Finally, the short video is automatically spliced ​​together according to the target duration. If the process is interrupted, it can be resumed through a breakpoint resume mechanism, or the parameters can be configured to adapt to the live-stream editing needs of different product categories.

[0216] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An automated intelligent editing method for live streaming of clothing, characterized in that, Includes the following steps: The original video of the clothing live stream is obtained, and the audio stream of the original video is segmented according to a preset duration to obtain multiple audio segments. The intelligent agent is scheduled to perform automatic speech recognition on the audio segment to obtain sentence-level recognition results containing text content, start timestamp, and end timestamp, and the timestamp is mapped to the absolute timeline of the original video; The sentence-level recognition results are subjected to two-layer prohibited content filtering. The first layer is a procedural compliance filtering based on dictionary and pattern matching, and the second layer is semantic discrimination based on a large language model of the clothing product description agent. After filtering, the set of retained texts of clothing product descriptions is obtained. The retained text set is subjected to Agent-based intelligent extraction of highlighted selling points based on a large language model. After splitting the retained text into clauses according to punctuation, clothing selling point words in the original text are extracted and a set of highlighted words is formed by prioritizing long words and retaining non-overlapping words. Based on the absolute timestamps of the reserved text set, the original video is sliced ​​at frame level precision to obtain multiple video segments; Perform multi-boundary constraint splitting on the subtitle text corresponding to each video segment to ensure that the highlighted words remain intact and are not split. Perform visual rendering on the highlighted words in the subtitles and complete the subtitle burning process; Multiple video clips are grouped and spliced ​​according to the target duration of the short video to generate the final short video.

2. The automated intelligent editing method for live streaming of clothing as described in claim 1, characterized in that, The programmatic compliance filtering includes static dictionary matching and dynamic interaction pattern preliminary matching. The static dictionary matching performs text inclusion matching and removes matching text based on the violation word library, sensitive word library, promotion word library, interactive language library, and casual conversation word library. The dynamic interaction pattern preliminary matching identifies and removes simple interactive guidance language such as "debit + number" and "debit + Chinese number" through regular expressions. The accurate identification and removal of complex interactive guidance language and non-product explanation content is completed by the semantic discrimination of the clothing product explanation agent intelligent agent based on a large language model, which is the core protection link of text screening.

3. The automated intelligent editing method for live clothing editing according to claim 1, characterized in that, The semantic discrimination of the clothing product explanation agent based on the large language model adopts a structured batch processing mechanism, which encapsulates candidate text into structured data containing timestamp and text fields and inputs them into the model in batches; the model output retains the original text content and timestamp, and a backup retention strategy is enabled when the model call is abnormal, so as not to discard candidate text.

4. The automated intelligent editing method for live clothing editing according to claim 1, characterized in that, The Agent-based intelligent extraction of selling points based on a large language model only extracts words, idioms, or phrases that exist in the original text, without rewriting or generating the text; the extracted clothing selling point vocabulary categories include fabric material, pattern design, craftsmanship details, wearing scenarios, functional characteristics, and wearing experience.

5. The automated intelligent editing method for live streaming of clothing according to claim 1, characterized in that, The subtitle multi-boundary constraint splitting includes punctuation boundary segmentation, word segmentation boundary secondary segmentation, and highlighted word boundary correction. When the candidate segmentation point is located inside the highlighted word, the segmentation point is automatically drifted to the front or back boundary of the highlighted word to ensure the integrity of the highlighted word.

6. The automated intelligent editing method for live streaming of clothing according to claim 1, characterized in that, After the subtitles are split, the total display time of the subtitles is allocated based on the proportion of the number of characters in each subtitle segment to the total number of characters, so that the display rhythm of the subtitles is consistent with the density of the audio information.

7. The automated intelligent editing method for live streaming of clothing as described in claim 1, characterized in that, Differentiated visual rendering is performed on highlighted words in the subtitles, and dual-link subtitle output mode is supported; in the differentiated visual rendering, short highlighted words with a length not exceeding the preset number of characters are rendered using a highlight color combined with an enlarged font size, while long highlighted words with a length exceeding the preset number of characters are rendered using a highlight color combined with an underline or outline; the dual-link output mode includes ASS subtitle script rendering mode and transparent PNG layer overlay rendering mode.

8. The automated intelligent editing method for live streaming of clothing as described in claim 1, characterized in that, The two-layer prohibited content filtering adopts a hierarchical progressive filtering and parallel batch processing architecture; first, a coarse screening is completed through procedural filtering, and then a fine screening is completed through semantic discrimination by an intelligent agent based on a large language model to explain clothing products; candidate texts are processed in parallel in batches, and when the model call fails, it automatically switches to the procedural filtering result as a backup.

9. The automated intelligent editing method for live streaming of clothing according to claim 1, characterized in that, The frame-level precision slicing adopts a decoding and re-encoding slicing method; the video segments are automatically grouped and spliced ​​according to a preset target duration interval. If the cumulative duration does not reach the lower limit, the segments are forcibly merged; if the upper limit is reached, the segments are grouped and output to ensure that the final video length is compliant and the semantics are coherent.

10. The automated intelligent editing method for live clothing editing according to claim 1, characterized in that, It also includes steps for resuming interrupted processes and reusing intermediate results, as well as steps for generalizing and expanding scenarios. The steps for resuming interrupted processes and reusing intermediate results involve an intelligent agent recording the processing status and intermediate results of each stage in real time and generating file hash values. When the process is restarted after an interruption, the validity of the intermediate results is verified by the hash value, and the completed processing stages are automatically skipped. The steps for generalizing and expanding scenarios involve updating the rule lexicon, adjusting the prompt words of the large language model, and modifying the selling point category configuration to quickly adapt to the live video editing needs of different product categories.