News splitting control method and system based on multiple modes

Through the integration of multimodal deep semantic transformer and multimodal feature, the problem of inaccurate boundary judgment in news dismantling is solved, high-quality news dismantling fragment generation is achieved, and the accuracy and fragment quality of dismantling is improved.

CN120358390AActive Publication Date: 2025-07-22CHANGJIANG DRAGON NEW MEDIA CO LTD

Patent Information

Application Number
CN202510841331.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The existing news article-breaking method lacks multimodal deep semantic analysis and dynamic regulation mechanisms, resulting in inaccurate boundary judgment, poor semantic coherence and ineffective fragment quality.

Method used

A multimodal deep semantic transformer is used to split short video streams, combine visual, audio and text modal features, and deep semantic modeling is carried out through a multimodal deep semantic transformer, and a modal focus change monitoring mechanism and structural consistency verification mechanism are introduced to optimize the decomposition of boundaries.

Benefits of technology

It significantly improves the accuracy and real-timeness of news decomposition, ensures the semantic coherence and structural rationality of decomposition fragments, and improves multimodal semantic consistency and content quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358390A_ABST
    Figure CN120358390A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing and artificial intelligence, and discloses a multi-modal-based news splitting control method and system, and the method comprises the steps: receiving a news short video stream, and splitting the news short video stream into multi-modal data; generating a corresponding semantic representation based on the multi-modal data, fusing each modal feature to form a unified semantic representation, and generating an initial splitting boundary list; and performing multi-modal semantic consistency verification on each fragment corresponding to the initial segmentation boundary, optimizing the initial segmentation boundary in combination with a quality evaluation result of each fragment, and outputting a high-quality fragment. The accuracy and the real-time performance of strip splitting are obviously enhanced; a structure consistency verification mechanism is combined, so that the semantic coherence and the structure rationality of the strip splitting fragments are effectively guaranteed; and dynamic semantic drift regulation and control and fragment quality comprehensive evaluation guided by circulation consistency are further introduced, so that the multi-modal semantic consistency and content quality of the output fragments are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data processing and artificial intelligence, and specifically to a news strip control method and system based on multi-modal. Background Art

[0002] With the rapid development of short video platforms and intelligent media technologies, the production and distribution methods of news content have undergone profound changes. Traditional long-form news videos have gradually been replaced by short video forms that are more acceptable to users. News strip technology, which automatically splits a complete news video into several short video segments with complete semantics and clear themes, has become a core link in content operation.

[0003] Existing news strip methods mostly rely on rule-based shallow feature extraction means, mainly by detecting surface features such as video scene changes, silent segments, key frame changes, or speech pauses to determine the splitting boundaries of videos. However, such methods have significant limitations: First, surface features cannot effectively reflect the true semantic structure of news content, easily leading to mismatches between the strip boundaries and semantic turning points, resulting in problems such as semantic breaks or information omissions; Second, the strip method driven by a single modality ignores the synergistic relationship between multi-modal information such as audio, video, and text, and it is difficult to handle the boundary determination challenges brought by multi-modal information interaction in complex news scenarios; In addition, most existing technologies use static thresholds or fixed rules, lacking the ability to perceive the dynamic changes of context semantics, resulting in the strip results lacking flexibility and intelligence, and it is difficult to ensure the semantic integrity and audio-visual coordination of the output segments. Summary of the Invention

[0004] In view of the above problems, the present invention is proposed.

[0005] Therefore, the technical problem solved by the present invention is: In the existing news strip methods, due to the lack of multi-modal deep semantic analysis and dynamic regulation mechanisms, the problems of inaccurate boundary determination, poor semantic coherence, and ineffective guarantee of segment quality occur.

[0006] To solve the above technical problems, the present invention provides the following technical solution: A news strip control method based on multi-modal, including: receiving a news short video stream and splitting the news short video stream into multi-modal data;

[0007] Generating corresponding semantic representations based on the multi-modal data, fusing the features of each modality to form a unified semantic representation, and generating an initial strip boundary list;

[0008] Performing multi-modal semantic consistency verification on each segment corresponding to the initial segmentation boundary, and optimizing the initial segmentation boundary in combination with the quality evaluation results of each segment to output high-quality segments.

[0009] As a preferred solution of the multi-modal based news strip control method described in the present invention, where: the receiving of the news short video stream includes setting the sliding window length to H seconds and the sliding step to K seconds, satisfying 0 < K < H; and intercepting a news short video stream segment of H seconds every K seconds for preprocessing; where H is the sliding window length, representing the duration of each intercepted segment; K is the sliding step, representing the time interval between two adjacent interceptions;

[0010] The multi-modal data includes visual modality, audio modality and text modality data;

[0011] Extract video frames from the preprocessed news short video stream at a preset frame rate, perform normalization processing on the video frames to generate visual features ;

[0012] Extract audio modality data from the preprocessed news short video stream at a fixed sampling frequency, convert it into a Mel spectrogram to generate audio features , and call an automatic speech recognition model to generate a text transcription result when there is no caption in the news short video stream;

[0013] Based on the text transcription result, use a word segmentation model to generate a text embedding feature T;

[0014] At each time step, combine to generate the corresponding multi-modal feature vector ; where, represents the visual feature at time step , represents the audio feature at time step , represents the text embedding feature at time step ; preprocess the multi-modal feature vectors of all time steps and construct a multi-modal feature sequence ; represents the multi-modal feature vector corresponding to the th time step; represents the th time step, represents the time step number.

[0015] As a preferred solution of the multi-modal based news strip control method of the present invention, wherein: the unified semantic representation includes receiving the multi-modal feature sequence, including visual features, audio features, and text embedding features; respectively using linear mapping functions to transform the visual features, audio features, and text embedding features into a semantic space of the same dimension, and adding temporal position encodings to each modal feature vector; adding the linearly mapped visual features, audio features, and text embedding features to the corresponding temporal position encodings to obtain the multi-modal fusion input vector for each time step.

[0016] As a preferred solution of the multi-modal based news strip control method of the present invention, wherein: inputting the unified semantic representation into a multi-modal deep semantic transformer for deep semantic modeling to identify semantic changes; based on the identified semantic change boundaries, candidate time steps are generated to form an initial strip boundary list;

[0017] The multi-modal deep semantic transformer, by introducing a trainable gating factor, adjusts the fusion ratio of text embedding features, visual features, and audio features according to the context semantics to achieve adaptive deep fusion of multi-modal features; through a modal attention change monitoring mechanism, modal attention entropy is embedded in the cross-attention process to continuously perceive the change of modal dominance and dynamically trigger semantic mutation detection to optimize the determination of strip boundaries; according to a preset structural consistency verification mechanism, a local temporal semantic relationship graph is constructed based on the candidate time steps of the boundary, and graph structure modeling is performed in combination with edge weights, and the semantic coherence and structural rationality of the candidate time steps of the boundary are verified through a graph neural network.

[0018] As a preferred solution of the multi-modal based news strip control method of the present invention, wherein: the architecture of the multi-modal deep semantic transformer includes a backbone encoding layer, a gated cross-attention fusion layer, a modal attention change monitoring mechanism, and a structural consistency verification mechanism;

[0019] The backbone encoding layer is composed of a stack of Transformer encoders with a preset number of layers, which captures long-range dependencies in the time series based on the self-attention mechanism and extracts global semantic features;

[0020] One gated cross-attention fusion layer is inserted every a layers of Transformer encoders. Using the text modal features as query vectors, and the visual and audio modal features as keys and values, cross-modal attention calculations are performed, and based on the trainable gating factor dynamically adjusts the information injection ratio of each modal feature;

[0021] wherein, a represents the preset number of intervening layers; represents the trainable gating factor at the th time step.

[0022] As a preferred solution of the multi-modal based news strip control method of the present invention, wherein: the modality attention change monitoring mechanism is embedded in the gated cross-attention fusion layer, calculates the attention distribution of text, visual and audio modality features at each time step, and calculates the modality attention entropy based on the attention distribution , and within consecutive time steps, when it is detected that the change amplitude of the modality attention entropy exceeds a preset threshold , local semantic trajectory deviation detection is performed;

[0023] The structure consistency verification mechanism includes, after the local semantic trajectory deviation detection, constructing a local temporal semantic relationship graph based on the deep semantic representation sequence , performing graph structure modeling in combination with edge weights, and performing weighted feature aggregation and structure analysis through a two-layer lightweight graph neural network to calculate the structure break confidence of the candidate strip boundary , verifying the semantic coherence and structural rationality of the boundary region; wherein, represents a set of points, represents a set of edges, represents the th time step's structure break confidence.

[0024] As a preferred solution of the multi-modal based news strip control method of the present invention, wherein: performing the multi-modal semantic consistency includes, based on each segment corresponding to the initial segmentation boundary, constructing a cross-modal mapping relationship with the text modality as the core; respectively mapping the text features to the visual modality space and the audio modality space through a cross-modal encoder, and then mapping back to the text space from their respective modality spaces, and taking the difference between the original text features and the features after the reverse mapping as the cycle consistency loss value;

[0025] Adjust the semantic topic drift detection threshold according to the cycle consistency loss value , the formula is expressed as:

[0026]

[0027] Wherein, represents the basic drift threshold, is the regulation coefficient; represents the exponential function, represents the segment 's cycle consistency loss value; represents the jth strip segment; j represents the segment index, represents the segment 's semantic topic drift detection threshold;

[0028] Based on the deep semantic representation sequence of segments, calculate the cosine similarity between consecutive time points. When the cosine similarity drops by more than a set threshold, it is determined that the semantic coherence of the segment is insufficient, and the current splitting boundary is redetermined.

[0029] A news splitting control system based on multi-modal, wherein:

[0030] A data module, which receives a news short video stream and splits the news short video stream into multi-modal data;

[0031] An initial segmentation boundary module, which generates corresponding semantic representations based on multi-modal data, fuses the features of each modality to form a unified semantic representation, and generates an initial list of splitting boundaries;

[0032] An output module, which performs multi-modal semantic consistency verification on each segment corresponding to the initial segmentation boundary, and combines the quality evaluation results of each segment to optimize the initial segmentation boundary and output high-quality segments.

[0033] A computer device, including: a memory and a processor; the memory stores a computer program, and is characterized in that: when the processor executes the computer program, the steps of the method described in any one of the present inventions are implemented.

[0034] A computer-readable storage medium, on which a computer program is stored, and is characterized in that: when the computer program is executed by a processor, the steps of the method described in any one of the present inventions are implemented.

[0035] Advantages of the present invention: The multi-modal news splitting control method provided by the present invention realizes the adaptive deep fusion of video, audio and text features by introducing a multi-modal deep semantic transformer, and improves the semantic modeling accuracy and context awareness ability of boundary determination in the news splitting process; through the modal attention change monitoring mechanism and dynamically triggered semantic mutation detection, the problem of incorrect boundary judgment caused by traditional static rules is avoided, and the accuracy and real-time performance of splitting are significantly enhanced; combined with the structural consistency verification mechanism, the semantic coherence and structural rationality of the split segments are effectively guaranteed; further introducing dynamic semantic drift regulation guided by cycle consistency and comprehensive evaluation of segment quality, the multi-modal semantic consistency and content quality of the output segments are improved. Description of the Drawings

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1Overall flowchart of the multi-modal news strip control method provided for the first embodiment of the present invention. Detailed implementation manners

[0038] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] Embodiment 1, referring to Figure 1 , an embodiment of the present invention, provides a multi-modal news strip control method, including:

[0040] S1: Receive a news short video stream and split the news short video stream into multi-modal data.

[0041] Receive the news short video stream through the streaming media protocol, and perform dynamic processing on the video stream using the sliding window mechanism. Set the sliding window length to H seconds and the sliding step to K seconds, where 0 < K < H; and intercept a news short video stream segment of H seconds every K seconds for preprocessing; where H is the sliding window length, indicating the duration of each intercepted segment; K is the sliding step, indicating the time interval between two adjacent interceptions; realize the streaming processing of the continuous video stream, reducing system latency and memory pressure.

[0042] The multi-modal data includes visual modality, audio modality, and text modality data.

[0043] For the accessed video stream, the system synchronously extracts visual modality, audio modality, and text modality data.

[0044] The extraction of visual modality data includes extracting video frames from the preprocessed news short video stream at a preset frame rate Adjust each extracted frame image to 224×224 pixels uniformly; perform normalization processing on the pixel values and map them to the interval [0, 1].

[0045] Extract audio modality data, extract the video audio track within the corresponding time period from the preprocessed news short video stream at a fixed sampling frequency, and the fixed sampling frequency can be set to 16 kHz.

[0046] Use the short-time Fourier transform to convert the audio signal into a Mel spectrogram and extract audio features; when there is no subtitle information, call the automatic speech recognition (ASR) model to generate a text transcription result.

[0047] Extracting text modal data includes, when there is a subtitle track in the video, parsing and extracting the time-aligned subtitle text; performing BERT tokenization on the transcribed text or subtitle text to form a token sequence, adding [CLS] and [SEP] tokens to obtain text embedding features; restricting the length of the text sequence, and truncating the excess part using a sliding window.

[0048] At each time step within, a corresponding multi-modal feature vector is combined and generated:

[0049]

[0050] Among them, represents the visual feature at time step , represents the audio feature at time step , represents the text embedding feature at time step ; The repetition filling strategy is used to align the multi-modal data in time steps to ensure that each time step contains complete tri-modal features, avoiding fusion bias caused by missing. After completing the time alignment, to improve the stability of the feature data and the model processing efficiency, the system performs mean-variance normalization processing on the visual feature and the audio feature respectively, normalizing the feature values to the standard normal distribution interval to eliminate the influence of different modal feature scale differences. The formula is as follows:

[0051]

[0052] Among them, represents the feature mean, represents the feature standard deviation, is the original feature vector.

[0053] At the same time, the text modality is mapped to a vector representation of a fixed dimension through the embedding layer to ensure a balance between dimension consistency and semantic expression ability. A standardized and time-aligned multi-modal feature sequence is constructed within consecutive time steps. The overall structure is represented as:

[0054]

[0055] Among them, represents the standardized and time-aligned multi-modal feature sequence; represents the multi-modal feature vector corresponding to the th time step; represents the th time step, represents the time step number, Represents the total number of time steps.

[0056] Furthermore, aiming at the deficiencies in the single processing method of news videos, lack of multi-modal collaboration and real-time and efficient processing ability in the existing technology, a method for dynamic splitting and multi-modal feature synchronous extraction of news short videos based on a sliding window mechanism is proposed. By setting the sliding window length and step size, the streaming processing of continuous video streams is realized, significantly reducing the system latency and memory occupancy, and improving the real-time performance and resource utilization rate of large-scale news video processing. At the same time, the extraction method of multi-modal data is improved to support the synchronous acquisition of visual, audio and text modal features within the same time step, avoiding the feature alignment deviation problem caused by asynchronous modal processing in traditional methods.

[0057] Even further, after completing the multi-modal feature extraction, a time-step-based alignment strategy and mean-variance normalization processing are introduced to solve the heterogeneity problems of different modal features in scale and time sequence, ensuring high consistency and stability in feature fusion within each time step. Compared with the existing technology that only relies on static frame extraction or single-modal feature input, it effectively improves the integrity of multi-modal feature expression and the model processing efficiency, provides standardized and unified high-quality input for subsequent deep semantic analysis and boundary determination, and significantly enhances the accuracy, robustness and overall system performance in the news splitting process.

[0058] S2: Generate corresponding semantic representations based on multi-modal data, and fuse the features of each modality to form a unified semantic representation to determine the initial segmentation boundary.

[0059] Receive the standardized and time-aligned multi-modal feature sequence output in step S1:

[0060]

[0061] Among them, , , are visual features, audio features and text embedding features respectively. The system maps different modal features to the same-dimensional semantic space through linear mapping, and the formula is expressed as:

[0062]

[0063] And combined with the time-step position encoding Generate the fused input vector:

[0064]

[0065] Among them, represents the overall multi-modal feature sequence, represents the single multi-modal feature vector at the time step ​ represents the index of the i-th time step; represents the time step of the visual feature vector, represents the time step of the audio feature vector; represents the time step of the text embedding feature vector; represents the visual feature after linear mapping result; represents the audio feature feature vector after linear mapping; represents the text embedding feature feature vector after linear mapping; , , respectively represent the linear mapping weight matrices corresponding to the visual, audio, and text modalities, which are used to transform the original features of different modalities into a unified semantic space of the same dimension, facilitating subsequent feature fusion and deep processing. represents the multi-modal fusion input vector at time step .

[0066] The fused multi-modal input vector sequence is input into a multi-modal deep semantic transformer for processing. The multi-modal deep semantic transformer architecture includes a backbone encoding layer, a gated cross-attention fusion layer, a modality attention change monitoring mechanism, and a structural consistency verification mechanism.

[0067] The backbone encoding layer performs deep semantic modeling on the unified mapped multi-modal feature sequence through a 60-layer Transformer structure. Aiming at the problem of insufficient context capture ability of traditional shallow models in processing long sequences, it uses an enhanced self-attention mechanism to improve the perception ability of long-distance semantic dependency relationships in the time series, achieving accurate extraction and dynamic understanding of the global semantic features in the complex context of news short videos, and providing a more complete and coherent semantic basis for subsequent boundary determination.

[0068] In the backbone encoding process, a gated cross-attention fusion layer is introduced to dynamically adjust the information injection ratio of the text, visual, and audio modalities according to the context, realizing adaptive multi-modal deep fusion.

[0069] The modality attention change monitoring mechanism generates a perception signal reflecting the change of modality dominance by statistically analyzing the attention distribution dynamics of each modality in the cross-attention in real time, which is used to trigger semantic mutation detection and boundary determination.

[0070] By integrating a gated cross-attention fusion layer during the backbone encoding process, it is possible to dynamically adjust the injection ratio of different modality information according to the context semantics, solving the problem of over-reliance on noisy modalities or insufficient information utilization under the traditional fixed fusion strategy, and significantly improving the adaptability and semantic modeling accuracy of multi-modal deep fusion;

[0071] At the same time, by constructing a modality attention change monitoring mechanism to track the dynamic modality distribution in cross-attention in real time and generate a perception signal of modality dominance change, replacing the traditional passive threshold detection method, it achieves the active triggering of the boundary determination process and the optimization of context awareness, effectively improving the response ability to complex semantic turns and the robustness of boundary recognition during the news strip splitting process. To improve the structural rationality and semantic coherence of boundary determination, a structural consistency verification mechanism based on a graph neural network (GNN) is introduced in the candidate strip splitting boundary region.

[0072] The formula for the input multi-modal feature sequence is expressed as:

[0073]

[0074] The backbone encoding layer includes 60 layers of Transformer encoders with 32 heads of self-attention each for deep semantic modeling. Each layer executes a standard self-attention mechanism to capture the long-range dependencies within the time series and extract global semantic features. The formula is expressed as:

[0075]

[0076] Among them, represents the overall multi-modal fusion feature sequence; represents the single multi-modal fusion feature vector at time step ; represents the th time step, represents the time step number; represents the total number of time steps of the feature sequence. SelfAttention(·) represents the self-attention mechanism function, Q represents the query matrix, K represents the key matrix, and V represents the value matrix; represents the transpose of the key matrix K; softmax(·) represents the softmax normalization function; represents the dimension of the key vector.

[0077] In the backbone layer, a takes 5; every 5 layers of Transformer encoders, 1 layer of gated cross-attention fusion layer is inserted, for a total of 12 layers, which are used to dynamically fuse text, visual, and audio modality features.

[0078] Using the text feature as the query vector , with visual and audio features as keys Sum value , calculate the cross-modal attention output , which is expressed by the formula:

[0079]

[0080] Introduce the trainable gating factor at the th time step , to achieve dynamic fusion control:

[0081]

[0082] Among them, represents the fusion feature vector at the th time step; represents the dimension of the key vector

[0083] Within each cross-attention layer, the attention distribution of different modalities is statistically analyzed in real time, and a modality attention change monitoring mechanism is introduced to calculate the modality attention entropy :

[0084]

[0085] Among them, represents the attention proportion of modality at time step (visual or audio). represents the modality attention entropy at time step , which is used to quantify the attention distribution of different modalities (visual, audio) in the cross-attention mechanism, reflecting the balance or bias of the current multi-modal information fusion. The table continuously tracks the change amplitude of the modality attention entropy. When the following conditions are met:

[0086]

[0087] It indicates that during the current multi-modal information fusion process, the dominant right ratio between vision and audio has changed significantly, usually corresponding to the turning point of the news semantic structure or the switching of content focus.

[0088] Therefore, instead of performing a global traversal, the system locally locates near this time step, activates the semantic trajectory deviation analysis and structural consistency verification mechanism, accurately determines whether there is a real semantic break in this area, avoids invalid calculations, and improves real-time performance and response efficiency.

[0089] Among them, represents the modality attention entropy at the previous time step ; Represents the trigger threshold for the change in modal attention entropy, which is set empirically. When the change amplitude exceeds this threshold, the system determines that a significant change in modal dominance has occurred, triggering the local boundary detection process. Represents at time step the change amplitude of the modal attention entropy.

[0090] Multimodal input After being processed by 60 layers of backbone Transformer encoding layers and 12 layers of gated cross-attention fusion layers, the model outputs the final hidden state at each time step of the 60th layer, denoted as: Output the final hidden state, denoted as:

[0091]

[0092] And the set of hidden states at all time steps is called the deep semantic representation sequence:

[0093]

[0094] Among them, Represents the time step corresponding deep semantic feature vector. Represents d-dimensional real number space.

[0095] Regarding the deep semantic representation sequence as the temporal trajectory of news content in the semantic space, which is used to reflect the dynamic process of semantic evolution over time. Based on this semantic trajectory, calculate the direction offset between adjacent time steps to identify potential semantic mutation points. The specific formula is:

[0096]

[0097] When occurs, the current semantic flow direction changes significantly, which is determined as a potential topic switching node and recorded as a candidate split boundary, entering the structural consistency verification stage. Represents the semantic trajectory direction offset at time step , the similarity and angular relationship between. and Represents the Euclidean norm (L2 norm) of the vector. Represents the time step corresponding deep semantic feature vector; Represents the time step corresponding deep semantic feature vector. Represents the semantic trajectory offset threshold, which is set empirically, .

[0098] The structural consistency verification mechanism includes constructing a local temporal semantic relationship graph for the time window of the candidate split boundary selected by the semantic trajectory offset detection 。

[0099] Select the deep semantic representation vector within the window as the nodes of the graph, which is expressed by the formula:

[0100]

[0101] Among them, represents the central time step of the candidate splitting boundary, represents the point set.

[0102] For each pair of adjacent nodes and , define the edge , and assign a weight to this edge, which is expressed by the formula:

[0103]

[0104] Among them, represents the edge set. represents the edge weight between node and node , and the value range is between [-1, 1]; represents the deep semantic feature vector at time step ; represents the deep semantic feature vector at time step .

[0105] When , it is marked as strongly connected, reflecting a high semantic correlation degree between nodes, and a higher weight is assigned during the GNN feature aggregation process to strengthen the feature stability of the semantic continuous region. When , it is marked as weakly connected, reflecting a low semantic correlation degree between nodes. During the graph neural network feature aggregation process, the weight contribution of this edge is dynamically adjusted to reduce its influence on node feature update and suppress the ineffective information transmission between low-correlation nodes. The proportion of weak connections in the neighborhood of the candidate splitting boundary nodes is statistically calculated, and combined with the feature difference degree of the nodes, the structural fracture confidence is calculated to improve the accuracy and robustness of boundary determination.

[0106] Input the local temporal semantic relationship graph into a two-layer lightweight GNN model, perform weighted feature aggregation based on edge weights, and update the node feature representation, which is expressed by the formula:

[0107]

[0108] Among them, represents the normalized edge weight:

[0109]

[0110] Among them, is the weight matrix of the l-th layer, represents the activation function; represents the node 's initial node feature. represents at the -th layer of the GNN, the updated feature vector of the node ; represents at the -th layer of the GNN, the feature vector of the neighbor node , as the input for feature aggregation. represents the set of neighbor nodes of the node ; represents the trainable weight matrix of the -th layer of the GNN, used to linearly transform the neighbor node features and extract high-order semantic relationship features; represents the node and its neighbor node 's normalized edge weight, represents the traversal variable of the neighbor node index, used for the summation operation in the denominator of the normalization to ensure that the sum of the edge weights of all neighbor nodes is 1, forming a weighted average mechanism.

[0111] After the GNN feature update is completed, perform a structural consistency analysis on the candidate strip boundary nodes, and calculate the feature difference degree between the candidate nodes and the adjacent nodes:

[0112]

[0113] Statistical average edge weight of the node :

[0114]

[0115] Comprehensively obtain the structural confidence score:

[0116]

[0117] Among them, represents at the time step the structural fracture confidence score calculated based on the graph neural network (GNN). represents the balance parameter of the structural confidence score, which is used to adjust the influence weights of the feature difference degree and the edge weight weakening on the overall score respectively, satisfying . represents the feature difference degree between the candidate strip boundary node and its adjacent time step nodes, represents the feature vector of the node after being updated by the GNN, represents the Euclidean distance of the vector. Represents the average edge weight of the node , and represents the edge weight between the node and its neighbor nodes .

[0118] When , it is determined that there is a significant semantic structure break at this candidate point, which is the initial segmentation boundary and enters the final multi-signal fusion scoring link to ensure the comprehensiveness and rigor of the boundary determination.

[0119] When , it is determined that there is no significant semantic structure break at this candidate point.

[0120] represents the structural break confidence threshold calculated based on the graph neural network (GNN), which is set according to requirements.

[0121] Integrate multi-dimensional decision signals and perform boundary confidence scoring

[0122]

[0123] When , then it is confirmed that this time step is a valid strip boundary and is included in the initial strip boundary list.

[0124] When , then it is confirmed that this time step is an invalid strip boundary.

[0125] Among them, represents the comprehensive boundary confidence score at time step , and the value range is [0, 1]; represents: the Sigmoid function; represents the semantic trajectory offset at time step ; represents the change amplitude of the modal attention entropy at time step ; represents the structural break confidence score obtained based on the structural consistency verification of the graph neural network (GNN) at time step ; represents the basic semantic difference degree at time step , which is defined as the average Euclidean distance between the deep semantic vectors of the current time step and its adjacent time steps. The formula is expressed as: ; represents the square operation on the basic semantic difference degree as a penalty term; represents the boundary confidence determination threshold, which is determined by optimizing the training set and is usually set in the interval.

[0126] The set of all boundary points that meet the valid strip-splitting boundaries is subjected to redundancy removal, and a non-maximum suppression strategy is adopted to eliminate duplicate boundaries caused by dense determination. Finally, a structured initial strip-splitting boundary list is output:

[0127]

[0128] Among them, represents the initial strip-splitting boundary list. represents the th time step in the time series, represents the specific index position of the boundary point in the time series, , and all confirmed boundary points are numbered in sequence, represents the total number of initial strip-splitting boundary points.

[0129] By constructing a unified multi-modal semantic representation mechanism, the accuracy and robustness of news strip-splitting boundary recognition are significantly improved. Different from the existing technologies that mostly adopt the methods of shallow feature splicing or simple weighted fusion, a linear mapping of multi-modal features, time position encoding enhancement, and deep semantic modeling mechanism are introduced. Specifically, the system first maps the visual, audio, and text embedding features to a unified semantic space to eliminate the differences in the original feature dimensions between modalities and the inconsistency in semantic expressions; then, combined with time position encoding, the context relationship in the time dimension is explicitly embedded into the fusion representation to obtain a multi-modal fusion input vector with sequence perception ability.

[0130] Furthermore, the above fusion vector is input into a deep Transformer encoder structure, and a gated cross-attention fusion layer is periodically introduced therein. The dynamic adjustment of the injection ratio of modal information under context awareness is realized through a trainable gating factor. Compared with traditional fusion strategies, this design has stronger modal adaptability and context semantic alignment ability, can effectively suppress the interference signals of redundant modalities, and improve the expression weight of key modalities.

[0131] In addition, through the calculation of modal attention entropy and the change monitoring mechanism, the traditional boundary detection method that relies on a fixed threshold strategy is replaced, which can actively trigger semantic mutation detection when the modal dominance fluctuates significantly, greatly improving the response ability of the strip-splitting system to semantic turning points. Combined with the structure consistency verification mechanism, this step can deeply analyze the structural coherence of the boundary candidate region through a graph neural network to ensure that the initial strip-splitting boundary breaks reasonably at the semantic change, overcoming the problem of insufficient utilization of structural information in existing solutions.

[0132] Furthermore, through the introduction of a number of innovative mechanisms such as multi-modal feature unified modeling, gated dynamic fusion, and structural coherence verification, the boundary recognition method has been systematically reconstructed from the feature layer, structural layer, and triggering mechanism layer, significantly improving the accuracy, robustness, and generalization ability of the news strip task in complex semantic scenarios, showing obvious technological progress.

[0133] S3: For each segment corresponding to the initial segmentation boundary, perform multi-modal semantic consistency verification, and optimize the initial segmentation boundary in combination with the quality evaluation results of each segment, and output high-quality segments.

[0134] According to the initial news strip boundary list , divide the news short video stream into several segments:

[0135]

[0136] For each segment , re-extract the standardized and time-aligned multi-modal feature sequence The formula is expressed as:

[0137]

[0138] For segment , construct a cross-modal mapping, calculate the cycle consistency loss, and the formula is expressed as:

[0139]

[0140] Among them, represents the cross-modal mapping function from text to vision, represents the reverse mapping from vision to text; represents the cross-modal mapping function from text to audio, represents the cross-modal mapping function from audio to text. represents the L2 norm. represents the j-th news strip segment; represents the total number of news strip segments. represents segment at time step the multi-modal feature vector; represents the time step index, with a range of , respectively represent the start and end time steps of segment . represents the cycle consistency loss value of segment , represents the overall representation of the text modality features of segment ; j represents the segment index.

[0141] Dynamically adjust the semantic drift detection threshold according to the cyclic consistency loss result :

[0142]

[0143] Among them, represents the basic drift threshold, is the regulation coefficient; represents the exponential function.

[0144] Based on the deep semantic representation sequence of the segment , calculate the semantic trajectory deviation degree:

[0145]

[0146] Among them, represents the deep semantic feature vector at time step ; represents the maximum deviation degree of the semantic trajectory of segment .

[0147] When , it is determined that there is a problem of insufficient semantic coherence in segment . Mark this segment as a semantic anomaly segment and trigger the boundary optimization mechanism to recalculate the segmentation boundary of this segment to avoid semantic chaos affecting the content quality of the short video. When , no operation is required.

[0148] Calculate the comprehensive quality score for segment , and the formula is expressed as:

[0149]

[0150] Among them, represents the effective information density, which is calculated as the ratio of the effective information duration to the total duration of the segment, reflecting the information richness of the segment. The higher the density, the more refined the content and the less redundancy. represents the audio-visual synchronization coordination, which calculates the semantic matching degree between the audio and video in the segment at each time node and statistics its fluctuation score. The higher the score, the more stable the audio-visual matching. represents the redundancy degree, which measures the information redundancy degree inside the segment by detecting the similarity of the deep semantic features in the continuous time period of the segment and statistics the proportion of the high repetition section. is the weight coefficient. represents the comprehensive quality score of segment .

[0151] When , it is marked as a low-quality segment. When , it is marked as a high-quality segment. Indicates the decision threshold for segment quality scoring, which is set according to requirements.

[0152] Integrate all high-quality segments and output a set of high-quality segments.

[0153] Among them, Indicates the th high-quality segment after optimization and confirmation; Indicates the number of segments after optimization.

[0154] Compared with the prior art, aiming at the problems of insufficient segment semantic coherence, poor modal coordination and inability to dynamically optimize segment quality in the news splitting process, a multi-modal semantic consistency optimization method integrating cyclic consistency verification and semantic drift regulation mechanism is proposed. By constructing a cross-modal bidirectional mapping with the text modality as the core, the cyclic consistency loss of segments is calculated in real time, realizing the quantitative determination of multi-modal semantic consistency, and breaking through the technical limitation of traditional segment evaluation that only relies on static rules or single-modal similarity. At the same time, based on the consistency loss, the semantic trajectory offset detection threshold is dynamically adjusted, enhancing the intelligent perception ability of internal semantic mutations and coherence anomalies in segments, and avoiding misjudgment or missed detection problems caused by fixed thresholds.

[0155] Based on the initial splitting boundary, multi-modal semantic consistency verification is performed on each divided segment, and combined with comprehensive quality evaluation, effectively improving the accuracy and expression quality of segment division. First, by constructing a bidirectional cross-modal mapping mechanism with the text modality as the core, semantic mapping relationships are established between text-vision and text-audio, and bidirectional cyclic consistency verification is respectively performed to quantify the semantic reconstruction errors in both directions, and the cyclic consistency loss value is obtained accordingly. Compared with traditional methods that only consider unidirectional alignment errors, this mechanism can comprehensively reflect the closed-loop consistency of semantic transfer between modalities, helping to identify potential semantic disconnection risks.

[0156] After obtaining the loss results, the system dynamically adjusts the decision threshold for semantic trajectory offset detection to adapt to the semantic stability degree of the current segment. Specifically, when the cyclic consistency loss is large, the threshold will be automatically tightened to enhance the sensitivity to semantic mutations; when the loss is small, the system relaxes the detection boundary to avoid misjudgment and enhance the robustness of the overall determination. This adaptive drift detection mechanism effectively overcomes the deficiencies of existing methods that rely on static thresholds and lack context awareness.

[0157] Furthermore, by combining the deep semantic representation of the segments, the evolution direction of the semantic vectors at each time step in the time series is analyzed to construct semantic trajectories, and the semantic deviation degree is calculated accordingly. If the deviation degree exceeds the adjusted drift threshold, the system can automatically identify the risk of internal semantic breaks in the segment and re-locate the segmentation boundary to ensure the semantic consistency and integrity within the segment. This mechanism makes up for the technical shortcomings of traditional methods in semantic coherence judgment, which rely on static structures and are prone to misclassification or missed judgment. Comprehensive quality assessment is performed on each segment, and weighted scores are given based on three dimensions: effective information density, audiovisual coordination, and information redundancy. This scoring system not only comprehensively covers the key indicators of short video segments in terms of information carrying, modality matching, and expression compactness, but also supports strategy customization in different usage scenarios through an adjustable weight mechanism. High-quality segments are retained and output, while low-quality segments are subjected to boundary optimization processing, effectively improving the structural rationality and dissemination value of the finally generated segments. The overall process reflects the deep integration of multi-modal closed-loop verification, threshold adaptive adjustment, and semantic structure reconstruction, significantly outperforming traditional boundary judgment methods based solely on modality feature mutations.

[0158] Embodiment 2 is an embodiment of the present invention, which provides a news strip splitting control system based on multi-modalities, including: a data module that receives a news short video stream and splits the news short video stream into multi-modal data.

[0159] An initial segmentation boundary module that generates corresponding semantic representations based on multi-modal data, fuses the features of each modality to form a unified semantic representation, and generates an initial news strip boundary list.

[0160] An output module that performs multi-modal semantic consistency verification on each segment corresponding to the initial segmentation boundary, and optimizes the initial segmentation boundary in combination with the quality assessment results of each segment, and outputs high-quality segments.

[0161] Embodiment 3 is an embodiment of the present invention, which is different from the previous two embodiments in that:

[0162] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0163] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.

[0164] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), fiber optic devices, and portable compact disc read-only memories (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0165] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0166] Embodiment 4, an embodiment of the present invention, provides a news strip control method and system based on multi-modal. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.

[0167] Select 10 long news videos released by a news media in 2024 as the test set. The average duration of the videos is 300 seconds, covering different types such as politics, finance, and sports, and the content is complex with frequent modal interactions, which is suitable for testing the stripping accuracy and fragment quality. The experimental environment is configured as: Intel Xeon processor, 32GB of memory, NVIDIA RTX 3080 graphics card, and the stripping control algorithm described in the present invention is deployed in the system.

[0168] First, set the sliding window length to 10 seconds and the step size to 5 seconds, receive the news short video stream and split it into three types of modal data: visual, audio, and text. For the visual modality, video frames are extracted at 1fps and normalized. The sampling rate of the audio modality is set to 16kHz and converted into a Mel spectrogram; for videos without subtitles, text transcripts are generated through an ASR model, and text embedding features are generated in combination with a word segmentation model to construct a standardized multi-modal feature sequence.

[0169] Subsequently, input the multi-modal feature sequence into a deep semantic transformer, relying on 60 layers of Transformer and a gated cross-attention fusion layer to complete deep semantic modeling and dynamic boundary determination. Through a modal attention entropy monitoring and structural consistency verification mechanism, an initial stripping boundary list is output. For each fragment, bidirectional cycle consistency verification and dynamic semantic drift regulation are performed, and in combination with three indicators: effective information density, audiovisual synchronization coordination, and redundancy, fragment quality evaluation and boundary optimization are completed, and finally high-quality news stripping fragments are output.

[0170] In the stripping experiments of 10 news videos, the method of the present invention generated a total of 82 stripping fragments. The core indicators are statistically as follows:

[0171] Average cycle consistency loss: 0.12.

[0172] Average semantic trajectory deviation: 0.35.

[0173] Average effective information density: 0.78.

[0174] Average audio-visual synchronization coordination score: 0.85.

[0175] Average redundancy: 0.08.

[0176] Proportion of high-quality segments: 93.90% (77 / 82).

[0177] Compared with the traditional rule-based strip method (the proportion of high-quality segments is about 76.50%), the present invention is superior in terms of segment quality, semantic coherence, and audio-visual coordination.

[0178] The experimental results show that the present invention shows significant advantages in the strip task of complex news videos. First, the average effective information density reaches 0.78, which is about 15% higher than that of the traditional method, indicating that the content of the output segments is more refined and the ineffective information is significantly reduced. The audio-visual synchronization coordination score remains at 0.85, verifying that the present invention can effectively guarantee the semantic synchronization of the commentary and the picture through the multi-modal deep fusion and dynamic regulation mechanism, avoiding the common problem of sound-picture disconnection in the prior art.

[0179] In addition, the redundancy is controlled at 0.08, which is significantly lower than 0.20 of the traditional method, indicating that the present invention can accurately identify and eliminate the repetitive information in the segments, improving the content compactness. Through the dynamic semantic drift regulation mechanism guided by cycle consistency, the potential semantic coherence problems are accurately identified, avoiding the output of semantic discontinuous segments, and ensuring the overall semantic integrity and structural rationality of the strip results.

[0180] Generally speaking, through the innovative multi-modal deep semantic transformer architecture, dynamic boundary determination mechanism, and segment quality optimization strategy, the present invention realizes the intelligent control of the news strip process, significantly improves the accuracy, real-time performance, and output segment quality of the strip, overcomes the deficiencies of the prior art in fixed rules, single-modal processing, and lack of dynamic optimization ability, has outstanding creativity and novelty, is applicable to various news short video automatic generation scenarios, and has good application prospects and promotion value.

[0181] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A news strip control method based on multi-modalities, characterized in that, Including: Receiving a news short video stream and splitting the news short video stream into multimodal data; Generating corresponding semantic representations based on the multimodal data, fusing the features of each modality to form a unified semantic representation, and generating an initial strip boundary list; For each segment corresponding to the initial segmentation boundary, performing multimodal semantic consistency verification, and combining the quality evaluation results of each segment to optimize the initial segmentation boundary and output high-quality segments.

2. The multi-modal based news strip control method according to claim 1, wherein: The receiving of the news short video stream includes setting the sliding window length to H seconds and the sliding step to K seconds, where 0 < K < H; and intercepting a news short video stream segment of H seconds every K seconds for preprocessing; where H is the sliding window length, representing the duration of each intercepted segment; and K is the sliding step, representing the time interval between two adjacent interceptions; The multimodal data includes visual modality, audio modality, and text modality data; Extract video frames from the preprocessed news short video stream according to a preset frame rate, perform normalization processing on the video frames, and generate visual features ; Extract audio modality data from the preprocessed news short video stream at a fixed sampling frequency, convert it into a Mel spectrogram, and generate audio features , and call an automatic speech recognition model to generate a text transcription result when there is no subtitle in the news short video stream; Based on the text transcription result, using a word segmentation model to generate text embedding features T; At each time step a corresponding multi-modal feature vector is combinatorially generated ; among which, represents the visual feature at time step , represents the audio feature at time step , represents the text embedding feature at time step ; the multi-modal feature vectors for all time steps are preprocessed and a multi-modal feature sequence is constructed; represents the multi-modal feature vector corresponding to the th time step; represents the th time step, represents the time step number.

3. The multimodal-based news strip control method according to claim 2, wherein: The unified semantic representation includes receiving the multimodal feature sequence, including visual features, audio features, and text embedding features; Respectively using a linear mapping function to transform the visual features, audio features, and text embedding features into a semantic space of the same dimension, and adding time position encoding to each modality feature vector; adding the linearly mapped visual features, audio features, and text embedding features and the corresponding time position encoding to obtain the multimodal fusion input vector at each time step.

4. The multimodal-based news strip control method according to claim 3, wherein: Inputting the unified semantic representation into a multimodal deep semantic transformer for deep semantic modeling to identify semantic changes; based on the identified semantic change boundary candidate time steps, generating an initial strip boundary list; The multimodal deep semantic transformer, by introducing a trainable gating factor, adjusts the fusion ratio of text embedding features, visual features, and audio features according to the context semantics to achieve adaptive deep fusion of multimodal features; through a modality attention change monitoring mechanism, embedding modality attention entropy in the cross-attention process to real-time perceive the change of modality dominance and dynamically trigger semantic mutation detection to optimize the determination of the strip boundary; according to a preset structure consistency verification mechanism, constructing a local temporal semantic relationship graph based on the boundary candidate time steps, combining edge weights for graph structure modeling, and verifying the semantic coherence and structural rationality of the boundary candidate time steps through a graph neural network.

5. The multi-modal-based news strip control method according to claim 4, wherein: The multimodal deep semantic transformer architecture includes a backbone encoding layer, a gating cross-attention fusion layer, a modality attention change monitoring mechanism, and a structure consistency verification mechanism; The backbone encoding layer is composed of a stack of Transformer encoders with a preset number of layers, capturing long-range dependencies in the time series based on the self-attention mechanism and extracting global semantic features; Insert one of the gated cross-attention fusion layers every a layers of the Transformer encoder, use the text modality features as the query vector, and the visual and audio modality features as the key and value to perform cross-modal attention calculation, and dynamically adjust the information injection ratio of each modality feature based on the trainable gating factor ; dynamically adjust the information injection ratio of each modality feature where a represents a preset number of spaced layers; represents the trainable gating factor at the 6. The multi-modal-based news strip control method according to claim 5, wherein: The modal attention change monitoring mechanism is embedded in the gated cross-attention fusion layer, calculates the attention distribution for text, visual, and audio modal features at each time step, and calculates the modal attention entropy based on the attention distribution. , and within consecutive time steps, when it detects that the change amplitude of the modal attention entropy exceeds a preset threshold , local semantic trajectory deviation detection is performed. The structure consistency verification mechanism includes, after the local semantic trajectory deviation detection, constructing a local temporal semantic relationship graph based on the deep semantic representation sequence , performing graph structure modeling by combining edge weights, and executing weighted feature aggregation and structure analysis through a two-layer lightweight graph neural network to calculate the structure break confidence of the candidate strip boundaries , and verifying the semantic coherence and structural rationality of the boundary region; where represents a set of points represents a set of edges represents the th structural break confidence at the time step 7. The multimodal-based news strip control method according to claim 6, wherein: Performing the multimodal semantic consistency includes, for each segment corresponding to the initial segmentation boundary, constructing a cross-modal mapping relationship with the text modality as the core; respectively mapping the text features to the visual modality space and the audio modality space through a cross-modal encoder, and then mapping back to the text space from their respective modality spaces, and using the difference between the original text features and the features after the reverse mapping as the cycle consistency loss value; Adjust the semantic topic drift detection threshold according to the cyclic consistency loss value , which is expressed by the formula as: ; Among them, represents the basic drift threshold, is the regulation coefficient; represents the exponential function, represents the segment of the cycle consistency loss value; represents the j-th split segment; j represents the segment index, represents the segment of the semantic topic drift detection threshold; Based on the deep semantic representation sequence of segments, calculate the cosine similarity between consecutive time points. When the decrease in cosine similarity exceeds a set threshold, it is determined that the semantic coherence of the segment is insufficient, and the current strip boundary is redefined.

8. A multimodal-based news strip control system using the method according to any one of claims 1-7, characterized in that: A data module that receives a news short video stream and splits the news short video stream into multimodal data; An initial segmentation boundary module that generates corresponding semantic representations based on multimodal data, fuses the features of each modality to form a unified semantic representation, and generates an initial strip boundary list; An output module that performs multimodal semantic consistency verification on each segment corresponding to the initial segmentation boundary, and optimizes the initial segmentation boundary in combination with the quality evaluation results of each segment to output high-quality segments.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the multimodal-based news strip control method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the multimodal-based news strip control method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Video content splitting method and device

    CN108965920A

  • Video splitting method and device

    CN113539304A

  • Intelligent news broadcasting system and method based on deep learning

    CN119854545A

  • Intelligent cataloging method for all-media news based on multi-modal information fusion understanding

    US20220270369A1

Cited By

  • Sky-ground comprehensive operation inspection method and system based on multi-source cooperation

    CN120974402A

  • Document analysis evaluation method and system based on multi-modal semantic consistency

    CN121052242A

  • Police sample video coarse segmentation method and law enforcement recorder video intelligent slicing method

    CN121661564A