Long video structured label generation method and system

Through multimodal analysis and semantic understanding, the problems of low efficiency, semantic deficiencies and redundancy in long video processing are solved, and high-quality structured tags are generated, which improves the efficiency of video content understanding and management, and realizes the creation of commercial value.

CN120340016AActive Publication Date: 2025-07-18海看网络科技(山东)股份有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510750037.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-18
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing long video processing technology has problems such as low efficiency, insufficient coverage, lack of semantics, redundancy of labels, insufficient space-time alignment capabilities, and lack of dynamic adjustment of label fusion weight allocation, making it difficult to generate high-quality structured labels.

Method used

Multimodal analysis and semantic understanding methods are adopted to generate text content with timestamps by obtaining long video files, using large language models to extract key event chains and shard suggestions, combining decision analysis to generate three-level tags, and dynamic fusion is carried out, including technical means such as OCR subtitle extraction, ASR speech transcription, BERT-Whitening feature enhancement, dynamic time regularization, SBERT text similarity calculation and knowledge graph expansion.

Benefits of technology

It significantly improves the quality of long video content understanding, generates high-precision structured labels, improves video content management and retrieval efficiency, enhances intelligent capabilities, and creates commercial value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340016A_ABST
    Figure CN120340016A_ABST
Patent Text Reader

Abstract

The invention discloses a long video structured label generation method and system, and mainly relates to the technical field of semantic recognition and processing. Comprising the following steps: acquiring a long video file, and performing data processing on the long video file to generate text content with a timestamp; based on the obtained text content, obtaining fragment suggestions through LLM analysis; performing decision analysis on the obtained fragment suggestion to generate a final fragment; and generating a third-level label for the generated final fragment, and dynamically fusing the third-level label to generate a structured label of the long video. The method has the beneficial effects that the problems of semantic deficiency, rough strip splitting, label redundancy and the like in long video processing are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic recognition and processing, and specifically to a method and system for generating structured tags for long videos based on multimodal analysis and semantic understanding. Background Art

[0002] With the continuous rapid development of the multimedia industry, a large amount of long video content is generated every day and distributed to major network audio-visual platforms. For these large numbers of long videos, their remarkable characteristics are complex video content, excessive duration, and diverse types, etc. For these long videos, to achieve high-efficiency and high-quality retrieval, recommendation, and management, a high-quality tag system is required as a support. Therefore, how to effectively and deeply understand the video content, analyze and summarize high-quality elements, and generate high-precision tags is particularly important.

[0003] However, traditional long video processing technologies have the following limitations: (1) Traditional long video tag generation mainly relies on manual annotation, or single-modal analysis, or simple superposition analysis of multiple modalities (such as only vision / speech), with low efficiency and insufficient coverage, resulting in missing key content or noise interference, and lacking spatio-temporal alignment ability; (2) Existing sharding technologies are mostly based on fixed time intervals or simple scene detection, lacking semantic coherence, and it is difficult to generate structured tags that reflect local and global semantics; (3) There are problems such as conflicts between OCR and ASR results, time axis misalignment, and text repetition in subtitle extraction, affecting the subsequent analysis effect; (4) There is a lack of a dynamic adjustment mechanism for weight distribution during multi-source tag fusion.

[0004] Therefore, there is an urgent need for a method and system for generating structured tags for long videos based on multimodal analysis and semantic understanding to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for generating structured tags for long videos, which effectively solves problems such as semantic loss, rough splitting, and tag redundancy in long video processing.

[0006] To achieve the above object, the present invention is realized through the following technical solutions: On the one hand, the present invention provides a method for generating structured tags for long videos, including the following steps: Step S1: Obtain a long video file, and perform data processing on the long video file to generate text content with time stamps; Step S2: Based on the plot semantic parsing of the LLM, use a large language model to extract key event chains, infer character relationships and scene types from the text content obtained in step S1, and give sharding suggestions in combination with the goal-oriented inference chain technology; Step S3: Conduct decision analysis on the sharding suggestions obtained in Step S2 to generate the final sharding; Step S4: Generate three-level labels for the final sharding generated in Step S3, and perform dynamic fusion on the three-level labels to generate structured labels for the long video.

[0007] Preferably, in Step S1, after performing data processing on the long video file, text content with timestamps is generated, including the following steps: Step S11: Perform multi-modal subtitle extraction on the long video file, and perform OCR subtitle extraction and ASR speech transcription, respectively obtaining an OCR subtitle sequence and an ASR subtitle sequence; The OCR subtitle extraction is specifically: integrating the PP-OCRv4 model to extract subtitle text in video key frames to obtain the OCR subtitle sequence: ; Among them, , represents the start and end times and text content of the th subtitle; The ASR speech transcription is specifically: integrating the FunASR model to transcribe speech into text words, generating dialogue text with timestamps, and obtaining the ASR subtitle sequence: ; Among them, , represents the start and end times and text content of the th subtitle; Step S12: Based on the improved BERT-Whitening feature enhancement algorithm and the dynamic time warping optimization algorithm, align the time series of the OCR subtitle sequence and the ASR subtitle sequence in Step S11; Step S13: For the unaligned parts in Step S12, perform semantic consistency verification, and judge whether to trigger the key frame arbitration mechanism based on the text similarity calculated by SBERT.

[0008] Preferably, Step S12 includes the following steps: Step S121: Perform data cleaning and standardization processing on the subtitle sequences output by ASR and OCR, and record the start and end times of each sentence; Step S122: Use the pre-trained BERT model to generate the original sentence vectors for each sentence. Input the text into the BERT model and extract the last layer of hidden states to obtain two feature sequences of OCR and ASR respectively: The OCR sequence set is: , where represents the sentence vector of the th OCR subtitle; The ASR sequence set is as follows: , where represents the sentence vector of the th ASR subtitle; Step S123: Calculate the covariance matrix for the sentence vector set , where is the vector dimension, is the number of samples, and the covariance matrix is: ; For the covariance matrix perform singular value decomposition , and construct a whitening matrix through the eigenvectors and eigenvalues after decomposition: , ; Multiply the original sentence vector by the whitening matrix to eliminate the correlation between dimensions and standardize the distribution, obtaining the whitened vector ; Perform principal component analysis on the whitened vector , select the first principal components, and obtain the vector after principal component analysis and whitening ; On the basis of principal component analysis and whitening, by left-multiplying the eigenvector matrix , obtain the vector after ZCA whitening: ; Perform L2 normalization on the vector after ZCA whitening to obtain the normalized vector , where is the L2 norm of the vector; Step S124: Combine text similarity and temporal deviation to define a composite distance function : ; Among them, represents the difference between the OCR and ASR sentences, represents the offset of the OCR and ASR temporal center points, represents the weight factor, ; Construct distance matrix , where ; Step S125: Define the cumulative distance matrix , and calculate the minimum cumulative path through dynamic programming: ; The operator selects the minimum cumulative cost from three adjacent paths: represents a path aligned vertically, represents a path aligned horizontally, represents a path aligned diagonally, where the boundary conditions are: , , ; Starting from the end point of the cumulative distance matrix backtrack the minimum cumulative distance path in reverse and record the optimal path, map the correspondence between the ASR and OCR sequences, and generate the aligned subtitle sequence.

[0009] Preferably, the step S13 is specifically: calculating the text similarity based on SBERT If , then trigger the key frame arbitration mechanism, including: Extract the video key frames within the conflict time period ; Call the image recognition model to generate a scene description ; Calculate and ; Select the subtitle with high similarity as the final result.

[0010] Preferably, the step S3 includes the following steps: Step S31: Extract the number of events per unit time and construct the event density feature: ; where represents the event occurrence rate per unit time, is the total number of events occurring during the observation period, is the total observation duration, and the ratio of the two constitutes the estimation form of the basic event rate parameter of the Poisson process. An exponential penalty term ‌ is introduced as the interval correction factor, is the average event interval time, and the exponential decay function maps the interval time to the interval (0,1]; When events occur frequently ( ), , it degenerates into the standard Poisson event rate; When events are sparse ( ), , reflecting the weight decay characteristic of low-frequency events; Step S32: Calculate the sentiment intensity based on the text sentiment analysis model: ; Among them, The average activation intensity reflecting the emotional arousal degree within the specified time window, Represents the frame processing and feature extraction process of the audio time series signal, Is the Audio waveform data of the time slice, Is an audio feature extraction network based on deep learning, Specifically refers to the arousal degree feature channel in the emotional dimension, Realize the calculation of the second-order statistics of the feature sequence, Is the total number of time window sampling points, and the absolute value operation is used to eliminate the feature polarity difference; Step S33: Detect the lens switching frequency through the HSV histogram difference: ; Among them, Represents the lens switching frequency per unit time, Is the total number of lens switches in the video, Is the total duration of the video, and this ratio reflects the original switching intensity. An adaptive duration adjustment factor Is introduced to realize the non-linear adjustment based on the video duration, Is the total duration of the video, and the constant term 5 sets the benchmark point of the duration threshold. The exponential decay function maps the duration to the interval (0,1). When , the adjustment factor quickly approaches 1; when , the adjustment factor shows an exponential decay trend; Step S34: Calculate the Q value for sharding decision-making: Define the state space: , among which, Is the event density of the current period, Is the sentiment intensity, Is the lens switching frequency, Is the cumulative duration that has not been sharded currently; Define the action space: ; Define the depth Network value function: ; Among them, , , , through And Realize high-order feature abstraction, Map the implicit features to the action value space, and the double ReLU activation layer introduces piecewise linear non-linearity to enhance the model's ability to fit complex state-action relationships; If , then output the sharding instruction and reset ; If , then continue to accumulate the current period duration.

[0011] Preferably, in step S4, the three-level labels include: basic labels, semantic labels, and derivative labels. The generation of the basic labels includes: Use the pre-trained CNN to extract frame-level features and obtain shard features through temporal max pooling: ; where is the video frame at time point , is the shard time range; Perform named entity recognition on the subtitle text to construct an entity frequency matrix: ; where is the text content of the th shard, is the rd class entity; Use SoundNet to extract voiceprint features and identify scene acoustic labels through spectral clustering: ; where is the frequency-divided audio, is the predefined class acoustic template feature vector; Generate basic labels: ; where , , are modal weights.

[0012] Preferably, the generation of the semantic labels includes: Use the LLM to parse the shard text to generate structured event triples: ; where is the head entity, is the tail entity, is the relationship type; Sentiment analysis model based on RoBERTa-large, calculate the shard sentiment vector: ; Among them, is the sentiment classification matrix, is the text content of the shard; Construct the spatio-temporal correlation matrix between shards: ; Among them, represents the and the similarity coefficient of two entities between shards ; Generate semantic labels: ; Among them, represents vector concatenation, represents extracting key event nodes, represents dynamic score aggregation.

[0013] Preferably, the generation of the derived label includes: Link the entity to Wikidata and extract n-hop associated entities: ; Among them, is the initial edge set, represents that the path length of the edge to does not exceed 2, that is, there is at most one intermediate vertex between the associated vertices of the two edges, is the target set, representing the edge set expanded from ; Use a two-channel CNN to detect visual-text metaphors: ; Among them, is the sigmoid function, , represents using a convolutional neural network to extract local and global semantic information of image features, represents using the bidirectional Transformer encoder of BERT to capture the context dependence of the text sequence; Cross-modal similarity calculation based on CLIP: ; Among them, is the predefined cultural symbol library, Indicates mapping visual input and text input into a joint embedding space using the CLIP model to generate cross-modal feature vectors. and text input to generate cross-modal feature vectors. Indicates encoding candidate elements to output vector representations in the same space. Indicates a preset similarity threshold obtained by calculating the cosine similarity between the joint feature vector and the candidate vector. Generate derivative labels: ; Retain the derivative labels with the first digits of confidence.

[0014] Preferably, in step S4, the dynamic fusion of the third-level labels to generate the structured labels of the long video includes: Define the video as a graph structure with nodes being shards and edge weights . The label propagation process is expressed as:

[0015] where is the label heat matrix at time , is the number of shards, is the number of labels, is the normalized adjacency matrix , is the heat retention coefficient; The final global label weight is expressed as: ; where is the time decay factor, is the start time of the shard; Design an exponential decay weight to reduce the long-tail effect of early shards: ; where the local confidence represents the local association strength of the -th shard to the -th label. Remove the labels with scores lower than the threshold and retain the labels; Sort according to the semantic label - derivative label - basic label hierarchy, and then sort the same-level labels in descending order of FinalScore to output the structured labels.

[0016] On the other hand, a system based on the long video structured label generation method as described above is provided, including: ​​​​​​​​​​Video acquisition and preprocessing module, for: obtaining a long video file, and generating timestamped text content after processing the data of the long video file; Segmentation suggestion analysis module, for: based on the plot semantic parsing of the LLM, using a large language model to extract key event chains, infer character relationships and scene types from the obtained text content, and giving segmentation suggestions in combination with the goal-oriented reasoning chain technology; Final segmentation generation module, for: making decision analysis on the obtained segmentation suggestions and generating the final segmentation; Structured label generation module, for: generating three-level labels for the generated final segmentation, and dynamically fusing the three-level labels to generate structured labels for the long video.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: The method and system for generating structured labels for long videos based on multimodal analysis and semantic understanding provided by the present invention can significantly improve the quality of long video content understanding and generate high-quality structured labels. Through multimodal collaboration, dynamic decision-making, and knowledge enhancement, breakthrough improvements are achieved in dimensions such as accuracy, efficiency, and intelligence, providing a systematic technical solution path for the field of video content understanding. This technical solution can be quickly embedded into existing video processing pipelines in various industries through modular capabilities, creating more commercial value while improving the efficiency of content value mining. Description of the Drawings

[0018] Figure 1 is the flowchart of the method of the present invention; Figure 2 is the schematic structural diagram of the system of the present invention. Detailed Embodiments

[0019] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.

[0020] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only relationship words determined for the convenience of describing the structural relationship of each component or element of the present invention, and do not specifically refer to any component or element of the present invention, and should not be construed as a limitation to the present invention.

[0021] Embodiment: Such as Figure 1As shown, this embodiment takes the 125-minute film and television drama "The Wandering Earth" as an example to provide a method for generating structured tags for a long video, including the following steps: Step S1: Get the target video file, input it into the subtitle extraction module, and generate text content with timestamp after data processing of the long video file: The OCR subtitles extracted based on the PP-OCRv4 model are as follows, where the timestamp is calculated based on the time point corresponding to the frame: OCR subtitles = [[63040,64040,"Quick, quick, quick"],[64040,66040,"Look, it's Jupiter"],[66040,67040,"The largest planet in the solar system"],[71040,72040,"Dad"],[72040,73040,"There is an eye on Jupiter"],[74040,75040,"That's a big storm on Jupiter"],[76040,77040,"What is a storm"],[78040,79040,"Jupiter is like a big balloon"],[80040,82040,"Ninety percent is hydrogen"],[83040, 85040,"Grandpa, what is hydrogen"],[88040,89040,"Dad drives the big rocket fuel"],[90040,91040,"Hydrogen is"],[91040,93040,"Dad drives the big rocket fuel"],[94040,95040,"Oh"],[95040,96040,"Liu Qi"],[96040,97040,"Yeah"],[97040,99040,"One day"],[99040,101040,"When you can see Jupiter without a telescope"],[102040,105040,"Dad will be back"],...]; The ASR subtitles extracted based on the FunASR model are as follows: ASR subtitles = [[61540,64575,"Quickly, quickly, look, it's Jupiter."],[64575,66975,"The largest planet in the solar system."],[66975,72705,"Dad, there's an eye on Jupiter."],[72705,75980,"That's a big storm on Jupiter."],[75980,76645,"What's that again?"],[76645,79500,"Jupiter is like a big balloon."],[79500,81885,"Ninety percent of it is hydrogen."],[81885,85925,"What is economy?"],[85925,89050,"Dad, it's the fuel for the big rocket."],[89050,94750,"Hydrogen is the fuel for the big rocket, Liu Qi."],[94750,97615,"One day..."],[97615,100885,"When you can see Jupiter without a telescope."],[100885,102785,"Dad will come back."],...]; Data cleaning and standardization processing are performed on the OCR subtitle and ASR subtitle texts, and the results are as follows: OCR subtitles = [[63040,64040,"Quickly"],[64040,66040,"Look, it's Jupiter"],[66040,67040,"The largest planet in the solar system"],[71040,72040,"Dad"],[72040,73040,"There's an eye on Jupiter"],[74040,75040,"That's a big storm on Jupiter"],[76040,77040,"What's that storm again"],[78040,79040,"Jupiter is like a big balloon"],[80040,82040,"Ninety percent of it is hydrogen"],[83040,85040,"Grandpa, what is hydrogen"],[88040,89040,"Dad, it's the fuel for the big rocket"],[90040,91040,"Hydrogen is"],[91040,93040,"Dad, it's the fuel for the big rocket"],[94040,95040,"Oh"],[95040,96040,"Liu Qi"],[96040,97040,"Mm"],[97040,99040,"One day"],[99040,101040,"When you can see Jupiter without a telescope"],[102040,105040,"Dad will come back"],...]; ASR subtitles = [[61540,64575,"Hurry up and look. It's Jupiter"],[64575,66975,"The largest planet in the solar system"],[66975,72705,"Dad, there's an eye on Jupiter"],[72705,75980,"That's a big storm on Jupiter"],[75980,76645,"What's that again"],[76645,79500,"Jupiter is like a big balloon"],[79500,81885,"Ninety percent of it is hydrogen"],[81885,85925,"What is the economy"],[85925,89050,"Dad, the fuel for the big rocket"],[89050,94750,"Dad, the fuel for the big rocket, Liu Qi"],[94750,97615,"Wait until one day"],[97615,100885,"When you can see Jupiter without a telescope"],[100885,102785,"Dad will come back"],...]; Then use a pre-trained BERT model (such as bert-base-chinese) to generate the original sentence vectors for each sentence. For example, for the subtitle "The largest planet in the solar system", the result is: { "start":66040, "end":67040, "text":"The largest planet in the solar system", "cls_vector":array([1.27813548e-01,-1.01697311e-01,-1.89450294e-01,...,-6.05204463e-01],dtype=float32) }, where cls_vector is a 768-dimensional vector; Construct two feature sequences for OCR and ASR respectively: , , where, represents the sentence vector of the th OCR subtitle , represents the sentence vector of the th ASR subtitle , is the number of OCR subtitle sentences, is the number of ASR subtitle sentences, For example, the OCR sentence vectors in this embodiment are: X=(array([[-0.21084408, 0.18425691, -0.8068817, ..., 0.21524063, -0.663201, -0.41471308], [0.1778533, -0.06230379, -0.29200676, ..., 0.5005508, -0.51869017, -0.5282888], [0.12781355, -0.10169731, -0.1894503, ..., -0.00722518, 0.42749327, -0.60520446], ..., [-0.1204564, 0.20805721, -0.00422067, ..., 0.0084686, 0.46712604, -0.37301672], [0.01117876, 0.38632497, 0.3542288, ..., 0.74560887, -0.0552956, -0.45850903], [-0.43215615, 0.00265869, -0.49723646, ..., 0.15225762, -0.0987256, -0.01773002]], shape=(1683, 768), dtype=float32),), The ASR sentence vector is: Y=(array([[0.30487418, -0.25255665, -0.8390225, ..., 0.56387115, -0.68002206, -0.39981446], [0.12781355, -0.10169731, -0.1894503, ..., -0.00722518, 0.42749327, -0.60520446], [0.20673126, -0.41274023, -0.03949136, ..., 0.6847591, 0.21732676, -0.26216733], ..., [0.19717567, -0.51321626, -0.85671175, ..., -0.11984089, -0.12622643, -0.06641263], [0.82386845, 0.03786982, -0.61562794, ..., 0.6586994, -0.20120588, -0.06905391], [0.13286236, 0.42928082, 0.07204984, ..., 0.31346616, -0.55816233, 0.084345]], shape=(1252, 768), dtype=float32),), Then, the OCR feature sequence and the ASR feature sequence are respectively subjected to BERT-Whitening whitening and PCA dimensionality reduction processing to obtain two standardized feature sequences of OCR and ASR: , , Among them, represents the sentence vector representation of the th OCR caption after PCA dimensionality reduction, represents the sentence vector representation of the th ASR caption after PCA dimensionality reduction.

[0022] For the standardized feature sequences and of OCR and ASR, calculate the cosine similarity and the temporal deviation between them respectively, and construct a composite distance matrix: , The composite distance matrix in this embodiment is: D=(array([[1.46150000e+02, 6.71900000e+02, 1.89140000e+03, ..., 2.22602400e+06, 2.22983100e+06, 2.23046475e+06], [5.96150000e+02, 2.21900000e+02, 1.44140000e+03, ..., 2.22557400e+06, 2.22938100e+06, 2.23001475e+06], [1.04615000e+03,2.30900000e+02,9.91400000e+02,..., 2.22512400e+06,2.22893100e+06,2.22956475e+06], ..., [2.11679615e+06,2.11598090e+06,2.11476140e+06,..., 1.09374000e+05,1.13181000e+05,1.13814750e+05], [2.11829615e+06,2.11748090e+06,2.11626140e+06,..., 1.07874000e+05,1.11681000e+05,1.12314750e+05], [2.13134475e+06,2.13052950e+06,2.12931000e+06,..., 9.48254000e+04,9.86324000e+04,9.92661500e+04]], shape=(1683,1252)),), Then, the dynamic time warping (DTW) algorithm is used to find the optimal path on the composite distance matrix to align the two sequences. The alignment result is corrected by filtering out low-similarity matches (with distances exceeding the threshold), handling one-to-many alignments (such as one ASR sentence corresponding to multiple OCR sentences), complementary correction, with the aligned subtitles being based on OCR, and complementary correction is performed using ASR subtitles while preserving punctuation marks, etc.; Semantic consistency verification is performed on the set of unaligned segments. The text similarity is calculated based on SBERT, and when the text similarity is lower than the threshold, the key frame arbitration mechanism is triggered.

[0023] For example, for the OCR subtitle "Grandpa, what is hydrogen?" and the ASR subtitle "What is the economy?" in this embodiment, since the text similarity between the two is 0.1577, which is far lower than the threshold (such as 0.7), the key frame arbitration mechanism is triggered. The key frames within the conflict time period [81885, 85925] are extracted, and the preset image recognition model is used to generate a scene description and extract the keywords "camping, chatting, night sky, tent, astronomical telescope, bonfire, car, balloon, hydrogen". The similarity between the OCR subtitle and the keyword data is calculated as 0.3286, and the similarity between the ASR subtitle and the keyword data is calculated as 0.059. The OCR subtitle with a higher similarity is selected as the final result; The final aligned subtitle data is as follows: aligned_subtitles=[[63040,66040,"Quickly, quickly look. It's Jupiter, the largest planet in the solar system."],[66040,67040,"The largest planet in the solar system."],[71040,73040,"Dad, there's an eye on Jupiter."],[74040,75040,"That's a big storm on Jupiter."],[76040,77040,"What's a storm?"],[78040,79040,"Jupiter is like a big balloon."],[80040,82040,"Ninety percent of it is hydrogen."],[83040,85040,"Grandpa, what's hydrogen?"],[88040,89040,"Dad, it's the fuel for the big rocket."],[90040,96040,"Hydrogen is the fuel for the big rocket, Liu Qi."],[97040,99040,"One day,"],[99040,101040,"when you can see Jupiter without a telescope,"],[102040,105040,"Dad will come back."],...].

[0024] Step S2: Based on the plot semantic analysis of the LLM, use the large language model to extract the key event chain, infer the character relationships and scene types from the text content obtained in Step S1, and give sharding suggestions in combination with the goal-oriented reasoning chain technology: For example, the key event chain (137 pieces) extracted in this embodiment is: keyword_events=[{"start_time":61540,"end_time":72705,"description": "Liu Qi discusses Jupiter's characteristics with his father, introducing Jupiter as the largest planet in the solar system and its Great Red Spot storm"}, {"start_time":72705,"end_time":81885,"description": "Explain the composition of Jupiter's atmosphere (90% hydrogen) and its relationship with rocket fuel"}, {"start_time":81885,"end_time":102785,"description": "Father promises to achieve family reunion through a space mission and signs the underground city lottery agreement"}, {"start_time":102785,"end_time":138815,"description": "The dying entrustment and emotional farewell of the father before an important space mission" {"start_time": 138815, "end_time": 196865, "description": "Reveal that the rapid aging of the sun leads to an earth crisis and initiate the Wandering Earth Plan"} {"start_time": 196865, "end_time": 246775, "description": "Details of the construction plan of the planetary engine and the functions of the navigator space station"},..., {"start_time": 7411200, "end_time": 7499470, "description": "The brightest star in the night sky and the hope of civilization"}]; For example, the potential sharding suggestions (23 pieces) extracted in this embodiment are as follows: clips = [{"start_time": 61540, "end_time": 246775, "title": "The Crisis of Jupiter and the Solar System"}, {"start_time": 246775, "end_time": 431425, "title": "The Farewell between Liu Qi and His Father"}, {"start_time": 431425, "end_time": 774270, "title": "The Initiation of the Wandering Earth Plan"},..., {"start_time": 7072235, "end_time": 7499470, "title": "The Brightest Star in the Night Sky"}];

[0025] Step S3: Conduct decision analysis on the sharding suggestions obtained in step S2 to generate the final shards: For example, the features extracted in this embodiment are as follows: Event density feature = [0.0000324, 0.0000217, 0.00000875,..., 0.0000173, 0.0000140], Emotional intensity = [0.45, 0.62, 0.78,..., 0.71, 0.82], Shot switching frequency = [0.345497732671129,0.324956672443674,0.382102438455256,...,0.309088883938516,0.2410822956652]; State space = [[0.0000324,0.45,0.345497732671129,427],[0.0000217,0.62,0.324956672443674,359],[0.00000875,0.78,0.382102438455256,325],...,[0.0000173,0.71,0.309088883938516,185],[0.0000140,0.82,0.2410822956652,185]]; The final slice boundary of the video is obtained by calculating the Q value as: [[61540,246775],[246775,431425],[431425,774270],...,[6713095,7072235],[7072235,7499470]].

[0026] Step S4: Generate three - level labels for the final shards generated in step S3, and perform dynamic fusion on the three - level labels to generate structured labels for the long video: 1) First, for all video shards, extract visual features, classification features, entity features, and audio features respectively, and fuse them to generate basic label data: For example, the features extracted in this embodiment are as follows: Visual feature labels = [["Science Fiction", "Disaster", "Family Ethics", "Political Game", "Social Experiment", "Space Opera", "Wasteland Punk", "Collectivism", "Spring Festival Culture", "Environmental Disaster"], ["Science Fiction", "Disaster", "Political Game", "Social Experiment", "Space Opera", "Environmental Disaster", "Spring Festival Culture", "Technological Ethics", "Collectivism", "Interstellar Migration"],..., ["Science Fiction", "Disaster", "Political Game", "Social Experiment", "Space Opera", "Environmental Disaster", "Technological Ethics", "Spring Festival Culture", "Collectivism", "Interstellar Migration"]], Classification feature labels = [["Jupiter", "Solar aging", "Underground city", "Lottery system", "Planetary engine", "Liu Qi", "Han Ziang", "Flint", "Roche limit", "Navigator"], ["Navigator", "Li Yiyi", "Flooding of Haikou", "Lottery system", "Sea level", "Planetary engine", "Spring Festival culture", "United Nations", "Astronaut", "Underground city planning", "United Government"],..., ["Earth", "Sun", "Liu Qi", "Han Ziang", "Han Duoduo", "Flint", "Planetary engine", "Navigator", "United Government", "Hibernation pod"]], Text entity feature labels = [["Jupiter storm", "Planetary engine", "Frozen city", "Rocket launch", "Tsunami flooding", "Data monitoring screen", "Spring Festival decoration", "Dialogue close-up", "Underground city entrance", "Earth orbit"], ["Space station", "Earth orbit", "Tsunami flooding", "Underground city entrance", "Spring Festival decoration", "Data monitoring screen", "Frozen city", "Rocket launch", "Government meeting room", "Rescue equipment"],..., ["Change of Earth orbit", "Melting of frozen city", "Dialogue close-up", "Data monitoring screen", "Spring Festival decoration"]], Audio feature labels = [["Crisis alarm", "Government broadcast", "Rocket countdown", "Natural wind and snow sound", "Electronic synthetic sound", "Festival music", "Mechanical operation sound", "Explosion sound", "Broadcast announcement", "Dialogue quarrel"], ["Space countdown", "News announcement", "Natural disaster sound effect", "Mechanical operation sound", "Festival music"],..., ["Global broadcast", "Dialogue", "Mechanical operation sound", "Festival music", "Electronic synthetic sound"]], The following basic label data is generated by integrating visual features, classification features, entity features, and audio features: Basic tags = [ (0.4, "Science Fiction", "Video Classification"), (0.4, "Disaster", "Video Classification"), (0.4, "Family Ethics", "Video Classification"), (0.4, "Political Game", "Video Classification"), (0.4, "Jupiter", "Entity Recognition"), (0.4, "Underground City", "Entity Recognition"), (0.4, "Planetary Engine", "Entity Recognition"), (0.4, "Spring Festival Culture", "Video Classification"), (0.4, "Frozen City", "Visual Feature"), (0.4, "Data Monitoring Screen", "Visual Feature") ], [ (0.4, "Science Fiction", "Video Classification"), (0.4, "Disaster", "Video Classification"), (0.4, "Political Game", "Video Classification"), (0.4, "Space Opera", "Video Classification"), (0.4, "Navigator", "Entity Recognition"), (0.4, "Sea Level", "Entity Recognition"), (0.4, "United Government", "Entity Recognition"), (0.4, "Space Station", "Visual Feature"), (0.4, "Rocket Launch", "Visual Feature"), (0.2, "Space Launch Countdown", "Audio Feature") ],..., [ (0.4, "Science Fiction", "Video Classification"), (0.4, "Disaster", "Video Classification"), (0.4, "Earth", "Entity Recognition"), (0.4, "Liu Qi", "Entity Recognition"), (0.4, "Earth Orbit Change", "Visual Feature"), (0.4, "Melting of Frozen City", "Visual Feature"), (0.2, "Global Broadcast Audio Feature"), (0.2, "Festival Music Audio Feature"), (0.4, "Spring Festival Reunion", "Visual Feature") ]]; 2) For all video segments, by constructing an event graph, analyzing, and analyzing the spatio-temporal correlation relationships of the segments, high-level semantic tags are inferred and obtained: For example, the features extracted in this embodiment are as follows: Event graph tags = [ ["Crisis Decision", "Intergenerational Commitment", "Technical Breakthrough", "Environmental Threat", "Resource Allocation", "Human Survival", "Institutional Game", "Civilization Continuity", "Task Execution", "Geopolitical Conflict"], ["Engineering Feat", "Environmental Threat", "Resource Allocation", "Human Collaboration", "Institutional Fairness", "Future Planning"], ..., ["Earth Survival", "Civilization Redemption", "Technical Dependence", "Rescue Operation", "Institutional Game"] ], Emotional intensity tags = [ ["Crisis Fear", "Family Warmth", "Rational Planning", "Sadness", "Hope"], ["Rational Planning", "Crisis Concern", "Hope", "Neutral", "Fear"], ..., ["Hope", "Warmth", "Rational Planning", "Sadness", "Neutral"] ], Space-time correlation tags = [["Solar expansion", "Jupiter's gravity", "Underground city evacuation", "Engine shutdown", "Roche limit", "Spring Festival crisis", "Nuclear fusion technology", "Ice escape", "Space station evacuation", "Rescue operation"], ["Long-term survival plan", "Solar expansion", "Sea level rise", "Space station evacuation", "Lottery system"],..., ["Earth push-off", "Jupiter's gravity", "Space station evacuation", "Flint transportation", "Roche limit"]], By integrating event graph analysis, sentiment intensity analysis, and space-time correlation analysis, the following semantic tag data is generated: Semantic tags = [(0.92, "Resource allocation", "Event graph"), (0.65, "Gravitational threat", "Space-time correlation"), (0.90, "Rational planning", "Sentiment intensity"), (0.88, "Earth's survival", "Event graph"), (0.85, "Spring Festival crisis", "Space-time correlation"), (0.80, "Institutional game", "Event graph"), (0.79, "Family ethical conflict", "Event graph"), (0.45, "Environmental disaster", "Space-time correlation"), (0.75, "Crisis fear", "Sentiment intensity"), (0.70, "Technical failure", "Event graph")], [(0.70, "Engineering feat", "Event graph"), (0.75, "Rational planning", "Sentiment intensity"), (0.28, "Environmental disaster", "Space-time correlation"), (0.79, "Human collaboration", "Event graph"), (0.80, "Institutional fairness", "Event graph"), (0.85, "Spring Festival crisis", "Space-time correlation"), (0.88, "Space station evacuation", "Event graph"), (0.90, "Neutral emotion", "Sentiment intensity"), (0.55, "Jupiter's gravity", "Space-time correlation"), (0.92, "Resource allocation", "Event graph")],..., [(0.92, "Civilization survival", "Event graph"), (0.65, "Hope inspiration", "Sentiment intensity"), (0.90, "Earth engineering", "Space-time correlation"), (0.88, "Hope psychology", "Event graph"), (0.85, "Disaster escape", "Event graph"), (0.80, "Global communication", "Space-time correlation"), (0.79, "Environmental threat", "Event graph"), (0.40, "Warm hope", "Sentiment intensity"), (0.75, "Starship civilization", "Space-time correlation"), (0.70, "Survival game", "Event graph")]; 3) For all video segments, through knowledge graph expansion, metaphor detection, and cultural correlation analysis, high-level semantic tags are inferred: For example, the features extracted in this embodiment are as follows: Knowledge Graph Expansion Tags = [["Roche Limit", "Nuclear Fusion Engine", "Oort Cloud Migration", "Climate Refugees", "Dyson Sphere Concept"], ["Oort Cloud Migration", "Climate Refugees", "Dyson Sphere Concept", "Starship Civilization", "Geoengineering"],..., ["Starship Civilization", "Climate Refugees", "Geoengineering", "Quantum Computer", "Hibernation Technology"]], Metaphor Detection Tags = [["Interstellar Escape Metaphor", "Myth of Sisyphus", "Noah's Ark", "Mechanical Romanticism", "Technological Ethical Dilemma"], ["Cosmic Engineering Metaphor", "Festival Crisis Narrative", "Technological Ethical Dilemma", "Wasteland Punk Aesthetics"],..., ["The Wandering Earth Metaphor", "Hope Motivation", "Mechanical Romanticism", "Technological Ethical Dilemma"]], Cultural Association Tags = [["Eastern Collectivism", "Spring Festival Reunion", "Family Reconstruction", "Confucian Ethics", "Disaster Narrative"], ["Festival Contrast", "Collectivism", "Craftsman Spirit", "Rural Complex", "Disaster Narrative"],..., ["Eastern Collectivism", "Disaster Narrative", "Thrifty Culture", "Collective Memory", "Spring Festival Reunion"]], Integrating knowledge graph expansion, metaphor detection analysis, and cultural association relationship analysis, the following derivative tag data is generated: Derivative Tags=[[(0.83, "Roche Limit", "Knowledge Graph Expansion"), (0.79, "Eastern Collectivism", "Cultural Association"), (0.80, "Interstellar Escape Metaphor", "Metaphor Detection"), (0.82, "Nuclear Fusion Technology", "Knowledge Graph Expansion"), (0.87, "Myth of Sisyphus", "Metaphor Detection"), (0.85, "Spring Festival Contrast", "Cultural Association"), (0.86, "Climate Refugees", "Knowledge Graph Expansion"), (0.88, "Technological and Ethical Dilemmas", "Metaphor Detection"), (0.90, "Planetary Engine", "Knowledge Graph Expansion"), (0.89, "Family Reconstruction", "Cultural Association")], [(0.90, "Eastern Collectivism", "Cultural Association"), (0.89, "Oort Cloud Migration", "Knowledge Graph Expansion"), (0.86, "Starship Civilization", "Knowledge Graph Expansion"), (0.88, "Space Opera Metaphor", "Metaphor Detection"), (0.82, "Dyson Sphere Concept", "Knowledge Graph Expansion"), (0.85, "Noah's Ark", "Metaphor Detection"), (0.87, "Spring Festival Reunion", "Cultural Association"), (0.79, "Geoengineering", "Knowledge Graph Expansion"), (0.80, "Doomsday Escape Metaphor", "Metaphor Detection"), (0.83, "Craftsman Spirit", "Cultural Association")],..., [(0.83, "The Wandering Earth Metaphor", "Knowledge Graph Expansion"), (0.80, "Spring Festival Reunion", "Cultural Association"), (0.79, "The Wandering Earth Metaphor", "Metaphor Detection"), (0.85, "Interstellar Migration", "Knowledge Graph Expansion"), (0.82, "Eastern Collectivism", "Cultural Association"), (0.87, "Nuclear Fusion Technology", "Knowledge Graph Expansion"), (0.86, "Interstellar Escape Metaphor", "Metaphor Detection"), (0.88, "Human Collaboration", "Knowledge Graph Expansion"), (0.90, "Mechanical Romanticism", "Metaphor Detection"), (0.89, "Thrifty Culture", "Cultural Association")]]; 4) Dynamically weighted fusion is performed on the three - level tag data of all shards to obtain the global tag results (Top10) as follows: Video Tags = [{"confidence": 0.92, "label": "Science Fiction"}, {"confidence": 0.9, "label": "Disaster"}, {"confidence": 0.88, "label": "Civilization Continuity"}, {"confidence": 0.85, "label": "Interstellar Escape Metaphor"}, {"confidence": 0.83, "label": "Environmental Disaster"}, {"confidence":0.80,"label":"Father-son bond"}, {"confidence":0.79,"label":"Eastern collectivism"}, {"confidence":0.76,"label":"Space engineering"}, {"confidence":0.74,"label":"Growth narrative"}, {"confidence":0.72,"label":"Technology dependence"}; Finally, after sorting, structured label data is output, as shown in Table 1.

[0027] Table 1 Structured label data output after sorting

[0028] As Figure 2 shown, this embodiment also provides a system based on the above long-video structured label generation method, including: Video acquisition and preprocessing module, used for: obtaining a long-video file, and generating timestamped text content after processing the long-video file; Segmentation suggestion analysis module, used for: based on the plot semantic parsing of the LLM, using the large language model to extract the key event chain, infer the character relationship and scene type of the obtained text content, and give segmentation suggestions in combination with the goal-oriented reasoning chain technology; Final segmentation generation module, used for: making decision analysis on the obtained segmentation suggestions to generate the final segmentation; Structured label generation module, used for: generating three-level labels for the generated final segmentation, and dynamically fusing the three-level labels to generate the structured labels of the long video.

[0029] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for generating long video structured tags, characterized in that It includes the following steps: Step S1: Obtain a long video file, and generate timestamped text content after processing the data of the long video file; Step S2: Based on the plot semantic parsing of the LLM, use the large language model to extract the key event chain, infer the character relationship and scene type from the text content obtained in Step S1, and give segmentation suggestions in combination with the goal-oriented inference chain technology; Step S3: Conduct decision analysis on the segmentation suggestions obtained in Step S2 to generate the final segmentation; Step S4: Generate three-level labels for the final segmentation generated in Step S3, and generate structured labels for the long video after dynamically fusing the three-level labels.

2. The method for generating long video structured tags according to claim 1, wherein In Step S1, after processing the data of the long video file to generate timestamped text content, it includes the following steps: Step S11: Perform multi-modal subtitle extraction on the long video file, and perform OCR subtitle extraction and ASR speech transcription to obtain the OCR subtitle sequence and the ASR subtitle sequence respectively; The OCR subtitle extraction is specifically: Integrate the PP-OCRv4 model to extract the subtitle text in the video key frames to obtain the OCR subtitle sequence: ; Among them, , indicating the start and end times and text content of the th subtitle; The ASR speech transcription is specifically: Integrate the FunASR model to transcribe the speech into text, generate the dialogue text with timestamps, and obtain the ASR subtitle sequence: ; Among them, , indicating the start and end times and text content of the th subtitle; Step S12: Align the time series of the OCR subtitle sequence and the ASR subtitle sequence in Step S11 based on the improved BERT-Whitening feature enhancement algorithm and the dynamic time warping optimization algorithm; Step S13: For the unaligned part in Step S12, perform semantic consistency verification, and judge whether to trigger the key frame arbitration mechanism based on the text similarity calculated by SBERT.

3. A method for generating long video structured tags according to claim 2, characterized in that, Step S12 includes the following steps: Step S121: Clean and standardize the subtitle sequences output by ASR and OCR, and record the start and end times of each sentence; Step S122: Use the pre-trained BERT model to generate the original sentence vectors for each sentence. Input the text into the BERT model and extract the last layer hidden states to obtain two feature sequences of OCR and ASR respectively: The OCR sequence set is as follows: , where represents the sentence vector of the th OCR subtitle; The ASR sequence set is as follows: , where represents the sentence vector of the th ASR subtitle; Step S123: For the sentence vector set calculate the covariance matrix, where is the vector dimension, is the number of samples, and the covariance matrix is: ; For the covariance matrix perform singular value decomposition , and construct a whitening matrix through the eigenvectors and eigenvalues after decomposition: , ; Multiply the original sentence vector by the whitening matrix to eliminate the correlation between dimensions and standardize the distribution, obtaining the whitened vector ; The whitened vector is subjected to principal component analysis, and the first principal components are selected to obtain the vector after principal component analysis whitening ; Based on the whitening of principal component analysis, by left-multiplying the eigenvector matrix , the vector after ZCA whitening is obtained: ; Perform L2 normalization on the vector after ZCA whitening to obtain the normalized vector , where is the L2 norm of the vector; Step S124: Define a composite distance function by combining text similarity and temporal deviation : ; Among them, represents the difference between the OCR and ASR sentences, represents the offset of the OCR and ASR time sequence center points, represents the weight factor, ; Construct distance matrix , where ; Step S125: Define the cumulative distance matrix , and calculate the minimum cumulative path through dynamic programming: ; The operator selects the minimum cumulative cost from three adjacent paths: Indicates the vertical alignment path, Indicates the horizontal alignment path, Indicates the diagonal alignment path, where the boundary conditions are: , , ; From the end point of the cumulative distance matrix Start, trace back the minimum cumulative distance path in reverse and record the optimal path, map the corresponding relationship between the two sequences of ASR and OCR, and generate the aligned subtitle sequence.

4. A method for generating long video structured tags according to claim 2, characterized in that The specific content of step S13 is as follows: calculating the text similarity based on SBERT If is satisfied, then trigger the key frame arbitration mechanism, including: Extract key frames of the video during the conflict time period ; Call the image recognition model to generate a scene description ; Calculation and ; Select the subtitle with high similarity as the final result.

5. A method for generating long video structured tags according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Extract the number of events per unit time and construct the event density feature; ; Among them, represents the event incidence rate per unit time, is the total number of events occurring during the observation period, is the total observation duration, and the ratio of the two constitutes the estimation form of the basic event rate parameter of the Poisson process. An exponential penalty term for the event interval ‌is introduced as the interval correction factor, is the average event interval time, and the exponential decay function maps the interval time to the interval (0, 1]; When events occur frequently ( ), , it degenerates to the standard Poisson event rate; When events are sparse ( ), it reflects the weight decay characteristic of low-frequency events; Step S32: Calculate the emotional intensity based on the text sentiment analysis model; ; Among them, The average activation intensity reflecting the emotional arousal degree within the specified time window, Represents the frame processing and feature extraction process of the audio time series signal, Is the Audio waveform data of the nth time slice, Is a deep learning-based audio feature extraction network, Specifically refers to the arousal degree feature channel in the emotional dimension, Realize the calculation of the second-order statistic of the feature sequence, Is the total number of sampling points in the time window, and the absolute value operation is used to eliminate the feature polarity difference; Step S33: Detect the shot transition frequency through the HSV histogram difference; ; Among them, represents the shot transition frequency per unit time, is the total number of shot transitions in the video, is the total duration of the video. This ratio reflects the original transition intensity. An adaptive duration adjustment factor is introduced to achieve non-linear adjustment based on the video duration. is the total duration of the video. The constant term 5 sets the benchmark point of the duration threshold. The exponential decay function maps the duration to the interval (0, 1). When , the adjustment factor quickly approaches 1; when , the adjustment factor shows an exponential decay trend; Step S34: Calculate the Q value for segmentation decision-making; Define the state space: , where is the event density in the current period, is the emotional intensity, is the shot transition frequency, is the cumulative duration of the current unfragmented part; Define the action space: ; Define depth Network value function: ; Among them, , , , through and achieve high-order feature abstraction, map the implicit features to the action value space, and the double ReLU activation layer introduces piecewise linear non-linearity to enhance the model's fitting ability for complex state-action relationships; If , output a sharding instruction and reset ; If , continue to accumulate the duration of the current period.

6. A method for generating long video structured tags according to claim 1, characterized in that In Step S4, the three-level labels include: basic labels, semantic labels, and derivative labels. The generation of the basic labels includes: Use the pre-trained CNN to extract the frame-level features, and obtain the segmentation features through temporal maximum pooling; ; wherein, is a time point of a video frame, is a sharding time range; Perform named entity recognition on the subtitle text to construct the entity frequency matrix; ; Among them, is the text content of the shard, is the class entity; Adopt SoundNet to extract the voiceprint features and identify the scene acoustic labels through spectral clustering; ; Among them, is the frequency-divided audio,[ is the predefined class of acoustic template feature vectors; Generate basic labels; ; Among them, , , are modal weights.

7. A method for generating long video structured tags according to claim 6, characterized in that, The generation of the semantic labels includes: Use the LLM to parse the segmented text to generate structured event triples: ; Among them, is the head entity, is the tail entity, is the relationship type; Based on the RoBERTa-large sentiment analysis model, calculate the shard sentiment vector: ; Among them, is the sentiment classification matrix, is the text content of the shard; Construct the spatio-temporal correlation matrix between shards: ; Among them, represents the between the similarity coefficient of two entities between slices; similarity coefficient; Generate semantic labels: ; Among them, represents vector concatenation, represents extracting key event nodes, represents dynamic score aggregation.

8. A method for generating long video structured tags according to claim 6, characterized in that, The generation of the derivative labels includes: Link the entity to Wikidata and extract n-hop associated entities: ; Among them, is the initial edge set, represents an edge to whose path length does not exceed 2, that is, there is at most one intermediate vertex between the associated vertices of the two edges. is the target set, representing the edge set expanded from ; Use a dual-channel CNN to detect visual-text metaphors: ; Among them, is the sigmoid function, , represents extracting local and global semantic information of image features using a convolutional neural network, represents using the bidirectional Transformer encoder of BERT to capture the contextual dependencies of the text sequence; Cross-modal similarity calculation based on CLIP: ; Among them, is a predefined cultural symbol library, represents mapping visual input and text input to a joint embedding space to generate cross-modal feature vectors, represents encoding candidate elements to output vector representations in the same space, represents a preset similarity threshold obtained by calculating the cosine similarity between the joint feature vector and the candidate vector; Generate derivative labels: ; Derivative tags before retaining confidence bits.

9. A method for generating long video structured tags according to claim 1, characterized in that In step S4, the dynamic fusion of the three-level labels to generate the structured label of the long video includes: Define the video as a graph structure , where the nodes are shards , and the edge weights . The propagation process of the labels is expressed as: Among them, is the label heat matrix at a moment, is the number of shards, is the number of labels, is the normalized adjacency matrix , is the heat retention coefficient; The final global label weight is expressed as: ; wherein is the time decay factor, is the starting time of the shard; Design an exponential decay weight to reduce the long-tail effect of early shards: ; Among them, the local confidence represents the local association strength of the th shard with respect to the th label. Remove the labels with scores lower than the threshold, and retain the labels; Sort according to the semantic label - derivative label - basic label hierarchy, and then sort the same-level labels in descending order of FinalScore, and output the structured label.

10. A system based on the long video structured tag generation method as described in claim 1, characterized in that, Include: Video acquisition and preprocessing module, used for: obtaining a long video file, and generating timestamped text content after data processing of the long video file; Shard suggestion analysis module, used for: based on the plot semantic parsing of the LLM, using the large language model to extract the key event chain, infer the character relationship and scene type from the obtained text content, and give shard suggestions in combination with the goal-oriented reasoning chain technology; Final shard generation module, used for: making decision analysis on the obtained shard suggestions and generating the final shards; Structured label generation module, used for: generating three-level labels for the generated final shards, and dynamically fusing the three-level labels to generate the structured label of the long video.

Citation Information

Patent Citations

  • Medical question and answer method and device based on large language model

    CN118410156A

  • Video work determination method and device, computer equipment, storage medium and program product

    CN119697404A

  • Full-automatic multi-mode machine labeling method for video with any length based on deep learning

    CN119763002A

  • Live broadcast subtitle generation method and related device

    CN119865669A

  • Long video tag generation method and device, electronic equipment and storage medium

    CN119988674A