A video shot structure generation method based on hierarchical semantic segmentation
By performing temporal segmentation preprocessing and cross-level semantic consistency verification on the video, a set of storyboard units is generated and structured tags are attached. This solves the problems in existing technologies that cannot identify semantic transitions without visual switching and lack a unified semantic standard for storyboard results, and achieves high-precision video storyboard generation and semantic retrieval adaptation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI XINTIANCE DIGITAL TECH CO LTD
- Filing Date
- 2026-04-28
- Publication Date
- 2026-07-31
AI Technical Summary
Existing video storyboard generation technology cannot recognize semantic transition scenes without visual switching, resulting in insufficient segmentation accuracy. The storyboard results lack a unified semantic standard, cannot guarantee the consistency between the storyboard content and the overall narrative logic of the video, and cannot be directly used for semantic retrieval and content analysis.
By performing temporal segmentation preprocessing on the input video, visual, audio, and subtitle text features are extracted, mapped to a preset video storyboard domain ontology semantic space, cross-level semantic consistency verification and dynamic determination of storyboard boundaries are performed, and a set of storyboard units is generated and attached with ontology semantic structured tags.
It achieves precise semantic segmentation, and the generated storyboard results have unified ontology semantic structured tags, which are suitable for semantic retrieval and content analysis, improving the semantic consistency and segmentation accuracy of the storyboard content.
Smart Images

Figure CN122493358A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of semantic processing and video structuring technology, and in particular to a method for generating video storyboard structures based on hierarchical semantic segmentation. Background Technology
[0002] Existing video storyboard generation technologies mostly segment based on inter-frame differences in low-level visual features, focusing only on pixel-level changes in the image. They fail to recognize semantic transitions without visual shifts, resulting in severely insufficient segmentation accuracy in special video content such as long dialogues and slow-motion narratives. Furthermore, existing layered segmentation schemes employ a unidirectional, progressive processing logic, causing errors from earlier steps to accumulate upwards, leading to global semantic gaps in the storyboard results and failing to guarantee consistency between the storyboard content and the overall narrative logic of the video. In addition, the storyboard results generated by existing solutions only contain temporal information and lack a unified semantic standard, making them unsuitable for direct use in downstream scenarios such as semantic retrieval and content analysis.
[0003] Based on the above problems, there is an urgent need for a video storyboard structure generation technology solution that can achieve accurate semantic segmentation and cross-level semantic consistency verification. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a video storyboard structure generation method based on hierarchical semantic segmentation, comprising the following steps: S1: Perform temporal segmentation preprocessing on the input video to obtain a set of basic temporal segments, and extract the visual features, audio features, and subtitle text features of each basic temporal segment in the set of basic temporal segments; S2: Map the visual features, audio features, and subtitle text features of each basic temporal segment to the preset video storyboard domain ontology semantic space to obtain the hierarchical semantic normalization features corresponding to each basic temporal segment; S3: Based on hierarchical semantic normalization features, cross-level semantic consistency verification and dynamic determination of storyboard boundaries are completed to obtain a set of video storyboard units; S4: Based on the video storyboard unit set and hierarchical semantic normalization features, generate ontology semantic structured tags corresponding to the storyboard units to complete the structured generation of video storyboards.
[0005] Preferably, the step of performing temporal segmentation preprocessing on the input video includes non-overlapping segmentation of the input video to obtain basic temporal segments, integrating all basic temporal segments into a set of basic temporal segments, and simultaneously extracting keyframe visual features, audio spectrum features, and subtitle text segmentation features for each basic temporal segment.
[0006] Further optimized, the preset video storyboard domain ontology semantic space is constructed based on the standard semantic tagging system of the video storyboard domain, and includes semantic mapping rules at four levels: frame-level semantic dimension, shot-level semantic dimension, scene-level semantic dimension, and narrative-level semantic dimension.
[0007] A further preferred step involves completing the cross-level semantic consistency verification and dynamic determination of the storyboard boundary, including: calculating the cross-level semantic similarity of adjacent basic time segments based on the hierarchical semantic normalization features of adjacent basic time segments; completing the preliminary determination of the storyboard boundary based on the cross-level semantic similarity; performing cross-level semantic consistency verification on the preliminary determined storyboard boundary; adjusting the position of the storyboard boundary; and integrating continuous basic time segments to obtain a set of video storyboard units.
[0008] A further optimized step, obtaining the hierarchical semantic normalized features corresponding to each basic time-series segment, is completed using the multimodal semantic ontology mapping confidence calculation formula. The multimodal semantic ontology mapping confidence calculation formula is as follows: ; In the formula, Modal numbering, Based on the basic time sequence segment numbering, The total number of dimensions in the preset video storyboard domain ontology semantic space. For the first The modality of the first The features of the 1st basic temporal segment are mapped to the 1st feature in the ontology semantic space. The dimensionless normalized eigenvalues of the dimension. For the pre-defined semantic space of video storyboard domain ontology, the first... The baseline weights of dimensional semantics, For the first The modality of the first The ontology mapping confidence of each basic time series segment is used to integrate the ontology mapping confidence of all modalities into hierarchical semantic normalized features of the corresponding basic time series segments.
[0009] A further preferred step of calculating the cross-level semantic similarity of adjacent basic time series segments includes calculating the ontology matching degree adjustment factor and semantic continuity adjustment factor for each modality based on ontology mapping confidence, completing the confidence weight allocation for each modality based on the ontology matching degree adjustment factor and semantic continuity adjustment factor, and completing the cross-level semantic similarity calculation of adjacent basic time series segments based on the confidence weight.
[0010] A further preferred step of completing the preliminary determination of the storyboard boundary includes calculating the confidence score of the storyboard boundary between adjacent basic temporal segments based on cross-level semantic similarity, and completing the preliminary determination of the storyboard boundary based on the confidence score of the storyboard boundary.
[0011] A further preferred step of generating ontology semantic structured tags corresponding to the storyboard unit includes extracting all hierarchical semantic normalization features corresponding to each video storyboard unit in the video storyboard unit set, matching corresponding frame-level semantic tags, shot-level semantic tags, scene-level semantic tags and narrative-level semantic tags for each video storyboard unit based on the preset semantic tag system of the video storyboard domain ontology semantic space, binding ontology semantic structured tags with the temporal information of the corresponding video storyboard unit, and generating video storyboard structured data.
[0012] A further preferred approach is to construct the pre-defined semantic space of the video storyboard domain, which includes sorting out all standard semantic tags in the video storyboard domain, constructing hierarchical relationships and semantic mapping rules between semantic tags, assigning corresponding dimensional weights to each semantic tag, and forming an ontology semantic space containing four levels of semantic dimensions.
[0013] A further preferred step is to extract keyframe visual features, audio spectrum features, and subtitle text segmentation features for each basic time-series segment. This includes extracting keyframes for each basic time-series segment, obtaining keyframe visual features through a visual feature extraction model, performing Fourier transform on the audio data of each basic time-series segment to obtain audio spectrum features, and performing word segmentation and vectorization processing on the subtitle text corresponding to each basic time-series segment to obtain subtitle text segmentation features.
[0014] Technical effects: This invention maps multimodal features to the ontology semantic space of video storyboard domain to generate hierarchical semantic normalization features. Combined with cross-level semantic consistency verification and dynamic determination of storyboard boundaries, it solves the core problems of existing technologies, such as the inability to identify semantic transitions without visual switching, accumulation of hierarchical segmentation errors, and global semantic discontinuity. The generated storyboard results have unified ontology semantic structured labels, which can be directly adapted to downstream semantic retrieval and content analysis scenarios. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the core process of a video storyboard structure generation method based on hierarchical semantic segmentation; Figure 2 A schematic diagram of the script's full-dimensional semantic deconstruction and weight mapping process; Figure 3 A schematic diagram of the process for quantifying narrative rhythm and adapting visual generation parameters; Figure 4 A schematic diagram of the overall process for hierarchical visual asset matching and standardized management; Figure 5 A schematic diagram of the incremental generation and hierarchical matching process for the visual asset feature library; Figure 6A schematic diagram of the standardized coding and storage iteration process for target multidimensional visual assets. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0017] Existing video segmentation technologies are mostly based on the inter-frame differences in visual features, focusing only on underlying visual changes and failing to identify semantic transitions without visual switching. The layered segmentation process suffers from error accumulation and global semantic discontinuity. The segmentation results lack a unified semantic standard and cannot be directly used for semantic retrieval and content analysis.
[0018] Based on this, please refer to Figures 1-6 This embodiment provides a video storyboard structure generation method based on hierarchical semantic segmentation. The complete technical solution is as follows: The input video undergoes temporal segmentation preprocessing to obtain a set of basic temporal segments. Visual features, audio features, and subtitle text features are extracted from each basic temporal segment. The temporal segmentation preprocessing steps include non-overlapping segmentation of the input video to obtain basic temporal segments, merging all basic temporal segments into a basic temporal segment set, and simultaneously extracting keyframe visual features, audio spectrum features, and subtitle text segmentation features from each basic temporal segment. The keyframe extraction steps for each basic temporal segment include extracting keyframes for each basic temporal segment, obtaining keyframe visual features using a visual feature extraction model, performing Fourier transform on the audio data of each basic temporal segment to obtain audio spectrum features, and performing word segmentation and vectorization processing on the corresponding subtitle text for each basic temporal segment to obtain subtitle text segmentation features. The input video sources include real-time captured video streams and locally stored video files. The video formats are compatible with mainstream encoding formats such as MP4, AVI, and MOV. Video resolutions support a full range from standard definition to 4K, and frame rates cover the standard 25fps to 60fps. During the non-overlapping segmentation of the input video, the total duration and frame rate information are first read, and the time interval between individual frames is calculated. Video frames are grouped based on a preset unit segmentation duration, which can be adapted according to the video type. Overlapping frames between adjacent segments are not retained during the segmentation process to avoid data redundancy and repeated calculations in subsequent feature extraction. The unit segmentation duration can be adaptively adjusted based on video type. For high-frame-rate, fast-changing video and short video content, a unit segmentation duration of 0.5 seconds is set to ensure the capture of rapidly changing semantic information. For moderate-frame-rate, slow-paced interview and documentary content, a unit segmentation duration of 1 second is set to reduce computation while maintaining semantic recognition accuracy. For low-frame-rate, slow-changing industrial monitoring and security content, a unit segmentation duration of 2 seconds is set to reduce unnecessary redundant calculations and improve processing efficiency. During segmentation, if the total video duration is not divisible by the unit segmentation duration, the remaining video content less than one unit segmentation duration is merged with the previous time segment to ensure that the duration of all basic time segments remains consistent and to avoid feature extraction deviations caused by differences in segment duration. After segmentation, each basic time sequence segment contains consecutive video frames, synchronized audio data, and corresponding subtitle text content. All basic time sequence segments are arranged in the order of video playback time and integrated to form a basic time sequence segment set. Each basic time sequence segment in the set is assigned a unique time sequence number for data location and association in subsequent processing.
[0019] Inter-frame difference calculations are performed on consecutive video frames within each basic temporal segment. By calculating the pixel grayscale value difference matrix between adjacent frames, the video frames corresponding to the peak inter-frame difference are selected as keyframes. Only one keyframe is retained for each basic temporal segment for subsequent visual feature extraction, significantly reducing the redundancy of visual feature calculations and improving processing efficiency. The visual feature extraction model employs a pre-trained feature extraction network based on a convolutional neural network. The network structure includes 8 convolutional layers and 4 pooling layers. The convolutional layers use 3×3 convolutional kernels with a stride of 1, and the pooling layers use 2×2 max-pooling kernels to extract spatial semantic features from the keyframe images. The output is a visual feature vector with a fixed dimension, and the dimension of the feature vector is consistent with the dimension of the subsequent ontology semantic space to ensure the adaptability of feature mapping. For each basic time-series segment, the audio data is processed by frame segmentation and windowing using a Hanning window function. The frame length is set to 25 milliseconds, and the frame shift is set to 10 milliseconds. A Fast Fourier Transform is performed on the segmented audio data to convert it to frequency domain data. Mel-spectral features are extracted to form audio spectral feature vectors. The dimension of these feature vectors is consistent with that of the visual feature vectors, achieving dimensionality uniformity across different modalities. For each basic time-series segment, the subtitle text content is preprocessed to remove punctuation and meaningless interjections. The preprocessed text is then subjected to Chinese word segmentation using a pre-trained language model-based algorithm. After segmentation, each word is vectorized to generate fixed-dimensional word vectors. These word vectors are then subjected to mean pooling to obtain the subtitle text segmentation feature vectors corresponding to the entire subtitle text content. The dimension of these feature vectors is consistent with that of the visual and audio feature vectors, completing the standardized extraction of features across the three modalities.
[0020] The pre-defined semantic space for the video storyboard domain is constructed based on the standard semantic tagging system of the video storyboard domain, and includes semantic mapping rules at four levels: frame-level semantic dimension, shot-level semantic dimension, scene-level semantic dimension, and narrative-level semantic dimension. The construction steps of the pre-defined semantic space for the video storyboard domain include: sorting out all standard semantic tags in the video storyboard domain, constructing hierarchical relationships and semantic mapping rules between semantic tags, and assigning corresponding dimensional weights to each semantic tag, forming an ontology semantic space containing four levels of semantic dimensions. The construction of the standard semantic tagging system for the video storyboard domain first involves sorting out all standard semantic tags in the video content analysis domain, covering four categories of tag content: main subject, action / behavior, scene environment, and narrative logic. Each category of tags is further divided according to a hierarchy from fine-grained to coarse-grained, forming a complete tag tree structure. The frame-level semantic dimension corresponds to fine-grained semantic information within the frame, including basic semantic tags such as the main subject, the main object, the subject's spatial position, the subject's action state, the color distribution of the frame, and the frame composition type. Each tag corresponds to a unique dimension index within the ontological semantic space, used to describe the core semantic content within a single frame. The shot-level semantic dimension corresponds to the shot semantic information of continuous frames, including shot movement type tags such as fixed shot, push-pull shot, panning shot, and focus shot; shot size tags such as wide shot, medium shot, close-up, and extreme close-up; and action continuity tags such as action start, action continuation, and action end, used to describe semantic changes within a single shot. The scene-level semantic dimension corresponds to the scene semantic information of shot combinations, including indoor scenes. The video includes tags for scenes, outdoor scenes, indoor scenes (detailed categories such as office, home, and industrial scenes), outdoor scenes (detailed categories such as street, natural, and architectural scenes), and event type tags such as dialogue events, operation events, and abnormal events within a scene. These tags are used to describe the semantics of a scene formed by the combination of multiple shots. The narrative-level semantic dimension corresponds to the global narrative information of the entire video, including narrative main line tags such as narrative beginning, narrative development, narrative climax, and narrative ending; narrative rhythm tags such as fast pace, slow pace, linear narrative, and flashback; and event causal relationship tags such as event cause, event process, and event result. These tags are used to describe the overall narrative logic of the video. The four dimensions of tags form a complete hierarchical semantic system, covering all semantic needs of video content analysis.
[0021] A hierarchical relationship is constructed between semantic tags, following a progressive relationship from frame-level semantic dimension to shot-level semantic dimension, shot-level semantic dimension to scene-level semantic dimension, and scene-level semantic dimension to narrative-level semantic dimension. This establishes a mapping relationship between tags, clarifies the aggregation rules from lower-level tags to higher-level tags, and the constraint rules from higher-level tags to lower-level tags, forming a coherent hierarchical relationship system. The semantic mapping rules define the feature vector mapping method for each semantic tag, assigning a unique dimension index to each semantic tag. The dimension indices of all semantic tags collectively constitute the total number of dimensions in the ontology semantic space, with each dimension corresponding to a unique semantic tag, ensuring a one-to-one correspondence of semantic information during feature mapping. A corresponding dimension weight is assigned to each semantic tag, with the weight value set based on the semantic tag's influence on scene segmentation. Semantic tags with a higher influence on scene boundary determination are assigned higher weight values. The weight values are fixed between 0 and 1, and the sum of the weight values of all dimensions is consistent with the total number of dimensions in the ontology semantic space, ensuring the rationality of the weight allocation. The completed ontology semantic space is stored in a dedicated embedded database. The database adopts a non-relational database structure, which supports fast query of semantic tags and real-time calling of dimension weights. It also supports the expansion and updating of the semantic tag system to adapt to the semantic segmentation needs of different video scenarios.
[0022] The visual features, audio features, and subtitle text features of each basic temporal segment are mapped to a predefined video storyboard domain ontology semantic space to obtain the hierarchical semantic normalized features corresponding to each basic temporal segment. The step of obtaining the hierarchical semantic normalized features corresponding to each basic temporal segment is completed using the multimodal semantic ontology mapping confidence calculation formula, which is as follows: ; The core theoretical basis of this formula is the cosine similarity calculation logic. It has been specifically adapted for ontology semantic space mapping scenarios. Its core function is to quantify the matching degree between single-modal features and the standard ontology semantic library, achieving unified semantic normalization of features from different modalities. This solves the core problem of inconsistent feature dimensions across modalities and the inability to perform cross-modal semantic comparison. In the formula… Each modality is assigned a unique, fixed number, corresponding to three different modality types: visual modality, audio modality, and subtitle text modality. This number is used to distinguish the computational data of different modalities. It is the basic time series segment number, which corresponds one-to-one with the unique time series number assigned in the basic time series segment set, and is used to locate the time series position corresponding to the computational data. The total number of dimensions of the preset video storyboard domain ontology semantic space is completely consistent with the total number of semantic tags in the ontology semantic space, ensuring that the calculation process covers all semantic dimensions. For the first The modality of the first The features of the 1st basic temporal segment are mapped to the 1st feature in the ontology semantic space. The dimensionless normalized eigenvalues of the dimension are obtained by mapping the original feature vector of the corresponding modality to the ontology semantic space through linear transformation, and then by the maximum and minimum value normalization process. The value range is fixed between 0 and 1 to ensure that all feature values involved in the calculation are dimensionless values, and to avoid the interference of dimensional differences on the calculation results. For the pre-defined semantic space of video storyboard domain ontology, the first... The baseline weight of semantic dimension is assigned during the construction of the ontology semantic space. It is a fixed dimensionless value with a range of 0 to 1. It is used to weight the feature values of different semantic dimensions to highlight the influence of the core semantic dimension on the matching results. For the first The modality of the first The ontology mapping confidence of each basic time segment is a dimensionless value with a fixed range between 0 and 1. The larger the value, the higher the semantic matching degree between the features of the modality and the corresponding dimension of the ontology semantic space, and the stronger the recognition of semantic information.
[0023] In the logical derivation of the formula, the numerator is the first... The modality of the first The dot product of the feature vectors of each basic temporal segment and the baseline weight vector of the ontology semantic space is calculated. The core logic of the dot product is to measure the co-directionality of the two vectors in the multidimensional space. The larger the dot product, the closer the directions of the two vectors are, and the higher the matching degree between the modal features and the ontology semantics. The denominator is the product of the L2 norms of the two vectors. The L2 norm is used to measure the magnitude of the vectors. The dot product of the magnitudes is normalized to eliminate the influence of the difference in vector length on the calculation result, ensuring that the confidence of the final output ontology mapping falls within a fixed interval of 0 to 1, so that the calculation results of different modalities and different temporal segments have a unified comparison standard. The dimensionless processing flow of this formula consists of two core steps. The first step is the linear transformation and normalization of the original feature vectors. The original feature vectors of each modality are first mapped to the corresponding dimension of the ontology semantic space through a linear transformation matrix. This linear transformation matrix is pre-trained based on the dimensional distribution of the ontology semantic space, ensuring a one-to-one correspondence between the transformed feature vectors and the ontology semantic dimensions. Then, a maximum-minimum normalization algorithm is used to compress the transformed feature values to the range of 0 to 1, eliminating the influence of differences in the numerical range of the original features from different modalities on the calculation results. The second step normalizes the numerator dot product result by multiplying the L2 norm of the denominator, ensuring that the final output confidence value is fixed within the range of 0 to 1. Regardless of the magnitude of the input feature vectors, the output results have a unified comparison standard, fully conforming to the principle of dimensional homogeneity. All parameters involved in the calculation are dimensionless values, the dimensions of addition and subtraction operations are completely unified, and the dimensional combinations of multiplication and division operations conform to mathematical rules, with no dimensional conflicts. In different types of video processing scenarios, the baseline weights of the ontology semantic space... Targeted adaptation is possible. For interview videos, the semantic dimension weight corresponding to dialogue content is set between 0.8 and 1, and the semantic dimension weight corresponding to scene changes is set between 0.2 and 0.4, highlighting the impact of semantic information of audio and subtitle text modal on scene segmentation. For film and television videos, the semantic dimension weight corresponding to shot changes and scene transitions is set between 0.7 and 0.9, and the semantic dimension weight corresponding to dialogue content is set between 0.3 and 0.5, highlighting the impact of semantic information of visual modal on scene segmentation. For industrial monitoring videos, the semantic dimension weight corresponding to abnormal events and action changes is set between 0.9 and 1, and the semantic dimension weight corresponding to static environmental information is set between 0.1 and 0.3, highlighting the recognition accuracy of abnormal semantic information and adapting to the scene segmentation requirements of different scenarios. The calculation process of the formula adopts a streaming computation method of time-series segments. After the feature extraction of each basic time-series segment is completed, the ontology mapping confidence calculation of that segment is immediately started. The calculation process is completed by the parallel computing unit of the graphics processor. The calculation latency of a single segment is controlled within 10 milliseconds to ensure the real-time performance of the entire processing flow. At the same time, the intermediate data and the final results generated during the calculation process are stored in the cache unit according to the time sequence number to provide data support for subsequent weight allocation and boundary determination. Based on the above calculation process and formula logic, those skilled in the art can complete the calculation of multimodal semantic ontology mapping confidence without creative labor, and realize the unified semantic normalization of features of different modalities.
[0024] Existing technologies typically use cosine similarity calculations only for comparing the similarity of two feature vectors with the same dimension. They do not adapt to multi-dimensional weighted scenarios in the ontology semantic space and cannot achieve normalized mapping of features from different modalities to a unified semantic space. This formula introduces a baseline weight vector from the ontology semantic space to weight the feature values of different semantic dimensions, achieving focused matching of core semantic dimensions and avoiding interference from irrelevant semantic dimensions. At the same time, through strict dimensionless processing, it ensures that the three different types of modal features—visual, audio, and subtitle text—can all be mapped to ontology semantics using a unified calculation formula, resulting in confidence values with a unified comparison standard. This solves the core problem of inconsistent feature dimensions across different modalities and the inability to compare semantic information across modalities, providing a unified semantic foundation for subsequent cross-modal semantic fusion and scene boundary determination. The ontology mapping confidence of all modalities is integrated according to the time sequence number and modality number to form the hierarchical semantic normalization feature of the corresponding basic time sequence segment. This feature contains four levels of semantic information and three modality matching confidence, which fully covers the full semantic content of the basic time sequence segment.
[0025] Based on hierarchical semantic normalization features, cross-level semantic consistency verification and dynamic determination of storyboard boundaries are performed to obtain a set of video storyboard units. The steps for performing cross-level semantic consistency verification and dynamic determination of storyboard boundaries include: calculating cross-level semantic similarity between adjacent basic time-series segments based on hierarchical semantic normalization features; performing preliminary determination of storyboard boundaries based on cross-level semantic similarity; performing cross-level semantic consistency verification on the preliminary determined storyboard boundaries; adjusting the position of the storyboard boundaries; and integrating continuous basic time-series segments to obtain a set of video storyboard units. The steps for calculating cross-level semantic similarity between adjacent basic time-series segments include: calculating the ontology matching degree adjustment factor and semantic continuity adjustment factor for each modality based on ontology mapping confidence; assigning credibility weights to each modality based on the ontology matching degree adjustment factor and semantic continuity adjustment factor; and calculating cross-level semantic similarity between adjacent basic time-series segments based on the credibility weights. The formula for calculating the dynamic weight allocation of modality credibility is: ; The core theoretical basis of this formula is the weight allocation logic of multi-attribute decision-making. It has been specifically modified for multimodal semantic credibility assessment scenarios. Through the collaborative constraints of ontology matching degree and semantic continuity, it achieves dynamic allocation of modal credibility. Its core function is to solve the problem of semantic credibility differences between different modalities in different video scenarios, ensuring that subsequent semantic similarity calculation results closely match the actual semantic changes in the video and avoiding calculation errors caused by fixed weight allocation. In the formula... For the first The modality of the first The ontology matching degree adjustment factor for each basic time segment is a dimensionless value with a fixed range between 0 and 1. This value is calculated based on the deviation between the ontology mapping confidence of the corresponding modality and the average ontology mapping confidence of historical time segments. The smaller the deviation, the higher the ontology matching stability of the modality. The larger the value of the adjustment factor, the stronger the positive impact on weight allocation. For the first The modality of the first The semantic continuity adjustment factor for each basic time segment is a dimensionless value with a fixed range between 0 and 1. This value is calculated based on the semantic change amplitude of the corresponding mode in the continuous time segment. The smaller the semantic change amplitude, the stronger the semantic continuity of the mode. The larger the value of the adjustment factor, the stronger the positive impact on the weight allocation. For the first The modality of the first The semantic continuity of a basic temporal segment relative to its predecessor is a dimensionless value, with a fixed range between 0 and 1. This value is obtained by calculating the cosine similarity of the ontology mapping confidence between the current temporal segment and the predecessor temporal segment of the corresponding modality. The larger the value, the stronger the semantic continuity between the two temporal segments. For the first The modality of the first The ontology mapping confidence of each basic time series segment is completely consistent with the value calculated by the preceding formula, ensuring the continuity and consistency of data during the calculation process. For the first The modality of the first The credibility weights of each basic time segment are dimensionless values, with a fixed range between 0 and 1. The sum of the credibility weights of all modalities is 1, which is used for weighting in the subsequent cross-level semantic similarity calculation process.
[0026] In the logical derivation of the formula, the numerator is the product of four values: ontology matching degree adjustment factor, ontology mapping confidence, semantic continuity adjustment factor, and semantic continuity. Through the synergistic effect of two adjustment factors and two core values, the semantic credibility of the corresponding modality in the current time segment is comprehensively evaluated. The larger the numerator value, the higher the semantic credibility of the modality, and the greater its weight in subsequent similarity calculations. The denominator is the sum of the numerator values of the three modalities. The summation result normalizes the numerator values to ensure that the sum of the credibility weights of the three modalities is 1, satisfying the mathematical rules of weighted calculation and avoiding interference from unbalanced weight allocation. The two-factor synergistic constraint logic of this formula consists of two core dimensions. The first dimension is the ontology matching stability constraint, through... Ontology mapping confidence Corrections are made when the confidence level of an ontology mapping for a particular mode fluctuates significantly across consecutive time segments. The value will automatically decrease, weakening the influence of this modality in weight allocation and avoiding weight allocation deviations caused by unstable modality matching; the second dimension is semantic continuity constraint, through... semantic continuity Corrections are made when a modality exhibits significant semantic changes and poor continuity across consecutive time segments. The value will automatically decrease, further weakening the weight of that modality and ensuring that the final weight truly reflects the semantic credibility of the corresponding modality. The calculation logic of the adjustment factor can be tailored to different video scenarios. For interview videos, the baseline value of the adjustment factor for the audio and subtitle text modality is set to 0.8, and the baseline value for the visual modality is set to 0.3, prioritizing the accuracy of dialogue semantic recognition. For film and television videos, the baseline value of the adjustment factor for the visual modality is set to 0.8, and the baseline value for the adjustment factor for the audio and subtitle text modality is set to 0.3, prioritizing the accuracy of shot and scene semantic recognition. The calculation process of the formula and the preceding ontology mapping confidence calculation adopt a streaming linkage method. After each time segment is completed... Calculate and immediately initiate the weight allocation calculation for this segment. During the calculation process, the following will be generated: , , Intermediate parameters are all stored in the cache unit to provide data support for subsequent boundary confidence calculation. Based on the above logic, those skilled in the art can complete the dynamic allocation of modal confidence weights without creative effort.
[0027] Existing technologies for multimodal feature fusion often employ fixed weight allocation, which cannot dynamically adjust based on the semantic credibility of different modalities in different scenarios. When a modality experiences semantic distortion due to environmental interference, fixed weights can lead to significant errors in the fusion result, directly affecting the accuracy of scene boundary determination. This formula achieves dynamic adaptive allocation of modal credibility weights through a two-factor collaborative constraint of ontology matching degree and semantic continuity. When the ontology matching stability of a modality decreases or its semantic continuity deteriorates, the corresponding weight will automatically decrease to reduce the impact of distorted modalities on the fusion result. Conversely, the weight ratio will increase to highlight the core role of high-credibility modalities. This solves the core problem that fixed weight allocation cannot adapt to the differences in semantic credibility across different scenarios, significantly improving the accuracy and robustness of cross-level semantic similarity calculation.
[0028] The initial determination of scene boundaries involves calculating the confidence score of scene boundaries between adjacent basic temporal segments based on cross-level semantic similarity, and then using this confidence score to make the initial determination of scene boundaries. The formula for calculating the confidence score of scene boundaries is: ; The core theoretical basis of this formula is the weighted cosine similarity calculation logic. It is specifically adapted for video scene boundary determination scenarios. Through dynamically allocated modal confidence weights, it weights and sums the cross-modal semantic similarities of adjacent time segments. Finally, the confidence score of the scene boundary is obtained by subtracting the weighted similarity from 1. Its core function is to quantify the degree of semantic difference between adjacent time segments, providing a precise quantitative basis for the initial determination of scene boundaries. This solves the core problem that fixed threshold determination cannot adapt to the semantic changes of different video content. In the formula... For the first The modality of the first The features of the 1st basic temporal segment are mapped to the 1st feature in the ontology semantic space. The dimensionless normalized eigenvalue of a dimension, the method of which is similar to... Completely identical, corresponding to adjacent subsequent basic time series segments, ensuring that the feature values of the two time series segments have the same dimensions and satisfy the mathematical rules for similarity calculation; For the first The modality of the first The credibility weights of each basic time segment are completely consistent with the values calculated by the preceding formula, and are used to weight the semantic similarity results of the corresponding modalities. The total number of dimensions of the preset video storyboard domain ontology semantic space is consistent with the total number of dimensions in the preceding formula to ensure that the calculation process covers all semantic dimensions. For the first The first basic timing segment and the first The confidence score of the storyboard boundary between basic time segments is a dimensionless value with a fixed range of 0 to 1. The larger the value, the greater the semantic difference between two adjacent time segments, and the more it conforms to the characteristics of the storyboard boundary.
[0029] In the logical derivation of the formula, the fractional part within the summation symbol is the first... In the mode, the The first basic timing segment and the first The cosine similarity of the feature vectors of the basic temporal segments is used to calculate the semantic similarity between adjacent temporal segments in a single modality. The larger the similarity value, the stronger the semantic continuity between the two temporal segments in that modality. The single-modal similarity results are weighted and summed using the modal credibility weights obtained from the previous calculation to obtain the cross-level semantic comprehensive similarity between adjacent temporal segments. The larger the comprehensive similarity value, the stronger the overall semantic continuity between the two temporal segments. By subtracting the cross-level semantic comprehensive similarity from 1, the confidence of the scene boundary is obtained, directly converting the semantic similarity into the boundary conformity, thus realizing the quantitative determination of the scene boundary. The weighted fusion logic of this formula fully inherits the dynamic weight allocation results from the previous step, ensuring that the modality with higher semantic credibility has a greater impact on the determination of the scene boundary. When there is no significant change in the visual modality, but a semantic transition occurs in the audio or subtitle text modality, the weight of the corresponding modality will automatically increase, ultimately reflected in a significant increase in the confidence of the scene boundary, achieving accurate recognition of semantic transition scenes without visual switching. The initial determination logic for scene boundaries is based on the changing trend of scene boundary confidence. When the confidence of the scene boundaries of three consecutive time segments continuously increases, and the confidence of the current segment exceeds the average confidence of historical segments, that position is determined to be the initial scene boundary. This replaces the fixed threshold determination method of existing technologies, and can adapt to the semantic change rhythm of different video content, avoiding false positives and false negatives caused by fixed thresholds. The calculation process of the formula is linked with the preceding weight allocation calculation in a streaming manner. After the weight calculation of each time segment is completed, the boundary confidence calculation of adjacent segments is immediately started. The calculation delay of a single segment is controlled within 5 milliseconds to ensure the real-time performance of the entire processing flow.
[0030] Existing technologies for determining scene boundaries often rely on setting a fixed threshold based on inter-frame differences in visual features. When the inter-frame difference exceeds the fixed threshold, it is determined to be a scene boundary. This approach cannot adapt to the semantic changes of different video content, nor can it identify semantic transition scenes without visual transitions, resulting in poor accuracy and adaptability. This formula uses dynamic weights to weight and fuse multimodal semantic similarity, integrating semantic information from visual, audio, and subtitle text modalities. Even if there are no significant changes in visual features, semantic transitions in the audio or subtitle text modal will be reflected in the changes in the confidence of the scene boundary, achieving accurate identification of semantic transition scenes without visual transitions. Furthermore, boundary determination is based on the changing trend of the confidence of the scene boundary, replacing the fixed threshold method of existing technologies. This adapts to the semantic change rhythm of different video content, solving the core problems of poor adaptability and inability to identify semantic transitions without visual transitions caused by fixed threshold determination, and significantly improving the accuracy and scene adaptability of scene boundary determination.
[0031] The initially determined storyboard boundaries undergo cross-level semantic consistency verification. The storyboard boundary positions are adjusted, and continuous basic temporal segments are integrated to obtain a set of video storyboard units. After the initial determination of storyboard boundaries, a preliminary set of storyboard boundary positions is obtained. For each initially determined storyboard boundary position, hierarchical semantic normalization features of a preset number of basic temporal segments before and after the boundary are extracted. The semantic consistency of adjacent storyboard segments is verified according to four semantic dimensions: frame level, shot level, scene level, and narrative level. During the verification process, the consistency between the higher-level narrative level semantics and scene level semantics is verified first, and then the consistency between the lower-level shot level semantics and frame level semantics is verified. When there is a consistency deviation in the higher-level semantics, but the lower-level semantics remain continuous, it is determined to be a semantic transition boundary, and the initially determined boundary position is retained. When there is a deviation in the lower-level semantics, but the higher-level semantics remain consistent, it is determined to be a misjudgment caused by local image changes, and the boundary position is removed. After verification, the reserved storyboard boundary positions are fine-tuned to ensure that the boundary positions are fully aligned with the semantic transition nodes of the video. According to the final determined storyboard boundary positions, the continuous basic time sequence segments are integrated to form independent video storyboard units. All video storyboard units are arranged in chronological order to form a video storyboard unit set. Each video storyboard unit is assigned a unique storyboard number for data association in the subsequent tag generation process.
[0032] Based on the video storyboard unit set and hierarchical semantic normalization features, ontology semantic structured tags corresponding to the storyboard units are generated, completing the structured generation of video storyboards. The steps for generating ontology semantic structured tags for storyboard units include: extracting all hierarchical semantic normalization features corresponding to each video storyboard unit in the video storyboard unit set; matching corresponding frame-level semantic tags, shot-level semantic tags, scene-level semantic tags, and narrative-level semantic tags for each video storyboard unit based on a pre-defined semantic tag system of the video storyboard domain ontology semantic space; binding the ontology semantic structured tags with the temporal information of the corresponding video storyboard unit to generate structured video storyboard data; extracting all hierarchical semantic normalization features corresponding to each video storyboard unit; integrating the hierarchical semantic normalization features of all basic temporal segments within the storyboard unit; and aggregating features according to four semantic dimensions to obtain the global semantic feature vector for each storyboard unit. Based on a pre-defined semantic tagging system within the ontology semantic space of the video storyboard domain, corresponding semantic tags are matched for each semantic dimension. During the matching process, the global semantic feature vector of the storyboard unit is compared with the tag feature vector within the ontology semantic space, and the semantic tag with the highest similarity is selected as the matching tag for the corresponding dimension, ensuring that the tag content completely corresponds to the semantic content of the storyboard unit. The ontology semantic tags at four levels are integrated to form an ontology semantic structured tag corresponding to the storyboard unit. The tag content contains all hierarchical semantic information of the storyboard unit. The ontology semantic structured tag is bound to the temporal information of the corresponding video storyboard unit. The temporal information includes the start time, end time, duration, and number of basic temporal segments contained in the storyboard unit. After binding, standardized video storyboard structured data is generated. The structured data is stored in a common JSON format and can be directly imported into downstream application platforms such as semantic retrieval systems, intelligent editing systems, and content analysis systems without secondary processing.
[0033] For a 30-minute interview video with a frame rate of 25fps and a resolution of 1920×1080, containing multiple topic transitions without video cuts, the method provided in this embodiment is used for processing. First, the input video is segmented without overlap, with each segment duration set to 1 second, resulting in 1800 basic temporal segments. Simultaneously, keyframe visual features, audio spectral features, and subtitle text segmentation features are extracted from each basic temporal segment. The feature vector dimension for each of the three modalities is set to 512 dimensions. The preset total number of dimensions in the video storyboard domain ontology semantic space is also considered. The value is set to 512, corresponding to 512 standard semantic tags. The four semantic dimensions each correspond to 128 semantic tags, and a corresponding baseline weight is assigned to each tag. Using the multimodal semantic ontology mapping confidence calculation formula, the ontology mapping confidence of each basic time segment across the three modalities is calculated. This integrates and forms hierarchical semantic normalization features. The confidence weights of the three modalities for each time segment are calculated using the modality confidence dynamic weight allocation formula. Then, using the confidence calculation formula for storyboard boundaries, the confidence scores of adjacent time segments are calculated. The initial determination of the storyboard boundaries was completed. Cross-level semantic consistency verification was performed on the initially determined boundaries to eliminate misjudged boundaries caused by local changes in the image, ultimately resulting in 24 video storyboard units, which completely correspond to the topic transition nodes within the video. Each storyboard unit was matched with four levels of ontology semantic structured tags, and after binding with temporal information, structured video storyboard data was generated, which can be directly used for semantic retrieval and topic classification of interview content.
[0034] The method provided in this embodiment achieves standardized comparison of semantic information of different modalities through unified ontology semantic mapping of multimodal features. Through dynamic weight allocation and cross-level semantic consistency verification, it solves the problems of error accumulation and global semantic discontinuity in existing technologies. Through multimodal fusion and segment boundary quantification judgment, it achieves accurate recognition of semantic transitions without visual switching. The generated structured segment data has a unified semantic standard and can be directly adapted to various downstream application scenarios.
[0035] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for video shot structure generation based on hierarchical semantic segmentation, characterized in that, Includes the following steps: S1: Perform temporal segmentation preprocessing on the input video to obtain a set of basic temporal segments, and extract the visual features, audio features, and subtitle text features of each basic temporal segment in the set of basic temporal segments; S2: Map the visual features, audio features, and subtitle text features of each basic temporal segment to the preset video storyboard domain ontology semantic space to obtain the hierarchical semantic normalization features corresponding to each basic temporal segment; S3: Based on hierarchical semantic normalization features, cross-level semantic consistency verification and dynamic determination of storyboard boundaries are completed to obtain a set of video storyboard units; S4: Based on the video storyboard unit set and hierarchical semantic normalization features, generate ontology semantic structured tags corresponding to the storyboard units to complete the structured generation of video storyboards.
2. The method of claim 1, wherein, The steps of performing temporal segmentation preprocessing on the input video include non-overlapping segmentation of the input video to obtain basic temporal segments, integrating all basic temporal segments into a set of basic temporal segments, and simultaneously extracting keyframe visual features, audio spectrum features, and subtitle text segmentation features for each basic temporal segment.
3. The method of claim 1, wherein, The pre-defined semantic space of the video storyboard domain is constructed based on the standard semantic tagging system of the video storyboard domain, and includes semantic mapping rules at four levels: frame-level semantic dimension, shot-level semantic dimension, scene-level semantic dimension, and narrative-level semantic dimension.
4. The method of claim 1, wherein, The steps to complete cross-level semantic consistency verification and dynamic determination of storyboard boundaries include: calculating cross-level semantic similarity of adjacent basic time segments based on hierarchical semantic normalization features; completing preliminary determination of storyboard boundaries based on cross-level semantic similarity; performing cross-level semantic consistency verification on the preliminary determined storyboard boundaries; adjusting the position of storyboard boundaries; and integrating continuous basic time segments to obtain a set of video storyboard units.
5. The method of claim 3, wherein, The step of obtaining the hierarchical semantic normalized features corresponding to each basic timing segment is completed by a multi-modal semantic ontology mapping confidence calculation formula, which is ; In the formula, Modal numbering, Based on the basic time sequence segment numbering, The total number of dimensions in the preset video storyboard domain ontology semantic space. For the first The modality of the first The features of the 1st basic temporal segment are mapped to the 1st feature in the ontology semantic space. The dimensionless normalized eigenvalues of the dimension. For the pre-defined semantic space of video storyboard domain ontology, the first... The baseline weights of dimensional semantics, For the first The modality of the first The ontology mapping confidence of each basic time series segment is used to integrate the ontology mapping confidence of all modalities into hierarchical semantic normalized features of the corresponding basic time series segments.
6. The video storyboard structure generation method based on hierarchical semantic segmentation according to claim 4, characterized in that, The steps for calculating the cross-level semantic similarity of adjacent basic time series segments include: calculating the ontology matching degree adjustment factor and semantic continuity adjustment factor for each modality based on ontology mapping confidence; completing the confidence weight allocation for each modality based on the ontology matching degree adjustment factor and semantic continuity adjustment factor; and completing the cross-level semantic similarity calculation of adjacent basic time series segments based on the confidence weight.
7. The method of claim 4, wherein, The steps to complete the preliminary determination of the storyboard boundary include calculating the confidence of the storyboard boundary between adjacent basic temporal segments based on cross-level semantic similarity, and completing the preliminary determination of the storyboard boundary based on the confidence of the storyboard boundary.
8. The video storyboard structure generation method based on hierarchical semantic segmentation according to claim 1, characterized in that, The steps for generating ontology semantic structured tags corresponding to storyboard units include extracting all hierarchical semantic normalization features corresponding to each video storyboard unit in the video storyboard unit set, matching corresponding frame-level semantic tags, shot-level semantic tags, scene-level semantic tags, and narrative-level semantic tags for each video storyboard unit based on the preset semantic tag system of the video storyboard domain ontology semantic space, binding ontology semantic structured tags with the temporal information of the corresponding video storyboard unit, and generating video storyboard structured data.
9. The method of claim 3, wherein, The pre-defined steps for constructing the ontology semantic space of the video storyboard domain include sorting out all standard semantic tags in the video storyboard domain, constructing hierarchical relationships and semantic mapping rules between semantic tags, assigning corresponding dimensional weights to each semantic tag, and forming an ontology semantic space containing four levels of semantic dimensions.
10. The method of claim 2, wherein, The steps for extracting keyframe visual features, audio spectrum features, and subtitle text segmentation features for each basic time-series segment include: extracting keyframes for each basic time-series segment; obtaining keyframe visual features through a visual feature extraction model; performing Fourier transform on the audio data of each basic time-series segment to obtain audio spectrum features; and performing word segmentation and vectorization processing on the subtitle text corresponding to each basic time-series segment to obtain subtitle text segmentation features.