Video training data generation method based on multi-modal semantic alignment
Through the multimodal semantic alignment video training data generation method, the problems of redundant data and missing key nodes in video analysis are solved, and efficient training data quality improvement and model performance improvement are achieved.
Patent Information
- Application Number
- CN202510861137.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies have problems with redundant data and missing key nodes in video analysis, resulting in waste of computing resources and reduced accuracy of analysis results.
A multimodal semantic alignment method is adopted to temporally align the audio, image frames and text information in the video, establish a cross-modal temporal mapping relationship, perform semantic enhancement processing, and dynamically grade the training samples according to semantic density and confidence to output structured training data.
It significantly improves the quality of training data, reduces redundant data, avoids missing key nodes, improves the accuracy and robustness of terminology recognition in professional fields, and reduces the cost of manual labeling.
Smart Images

Figure CN120766057A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio and video processing, and in particular to a method for generating video training data based on multimodal semantic alignment. Background Art
[0002] With the rapid development of information technology and the internet, video has gradually become one of the most important carriers of information dissemination in modern society. In particular, driven by emerging media forms such as short videos, self-media, and live streaming, the speed and scale of video data generation are growing exponentially. The ability to efficiently and accurately intelligently analyze and structure video content has become a key research topic in the fields of artificial intelligence and computer vision.
[0003] Traditional video analysis methods typically preprocess video data using fixed-interval or second-by-second sampling to generate training samples. For example, one or more frames are extracted every few seconds or at a fixed number of frames as training data. While this approach is effective for simple scenes and regular content, it often exhibits significant shortcomings in complex scenes or videos with uneven information distribution. This is because video content is inherently dynamic, and the semantic intensity, information density, and distribution of key content vary. Using fixed-interval frame extraction fails to reflect the actual importance of video content, leading to two typical problems: First, the generation of excessive redundant data—that is, excessive invalid segments in scenes with little or repetitive information, resulting in wasted computing resources; second, the loss of key nodes—that is, the inability to accurately and timely extract frames with high semantic value at nodes where content changes suddenly, semantic transitions occur, or important events occur, reducing the accuracy and reliability of the analysis results.
[0004] Therefore, how to improve the quality of training data from the source of video data and avoid the emergence of large amounts of redundant data and the loss of key nodes has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] The main purpose of the present invention is to provide a video training data generation method based on multimodal semantic alignment, aiming to improve the quality of training data from the source of video data and avoid the emergence of a large amount of redundant data and the loss of key nodes.
[0006] To achieve the above objectives, the present invention proposes a method for generating video training data based on multimodal semantic alignment, comprising the following steps: (1) Perform multimodal time alignment on the audio, image frames and text information in the video to establish a cross-modal temporal mapping relationship; (2) Performing semantic enhancement processing based on the time alignment results to improve the recognition accuracy of professional terms; (3) Dynamically classifying the training samples according to the semantic density and the confidence; (4) Outputting the classified structured training data to adapt to different training stages.
[0007] In an embodiment of the present application, the multi-modal time alignment in the step (1) comprises dynamic time window matching, and a tolerance range of the dynamic time window is adaptively adjusted according to at least one of a signal-to-noise ratio of an audio channel, a visual text density, and a time offset between speech and subtitles.
[0008] In an embodiment of the present application, the dynamic time window matching calculates a mapping relationship of an audio timestamp to a frame index by the following formula: F = τ * fps; Wherein, fps represents a video frame rate, and τ is a speech recognition output timestamp.
[0009] In an embodiment of the present application, the semantic enhancement processing in the step (2) comprises: Performing term replacement on the recognized text by a hierarchical term library, and the term library is activated according to a domain knowledge level; Resolving term ambiguity based on context semantic similarity, and a calculation formula is: ; Wherein, is an embedding vector of the term, is an embedding vector of a context window.
[0010] In an embodiment of the present application, the activation condition of the hierarchical term library comprises: L1-level terms are loaded by default in a specific video category; L2-level terms are activated when a word frequency exceeds a threshold or a context contains an associated keyword.
[0011] In an embodiment of the present application, the dynamic classification in the step (3) comprises: Primary training set: single-modal data and confidence higher than 0.85; Intermediate training set: cross-modal verification consistent and confidence between 0.70-0.85; Advanced training set: containing complex space-time relationship and confidence lower than 0.70, and requiring additional discriminant model labeling.
[0012] In an embodiment of the present application, the step (3) further comprises a training stage advancing mechanism, comprising: Switching stages when a verification set accuracy rate continuously improves by less than 1% for three training periods; At the end of the stage, adversarial samples are injected, including at least one of random mismatching of speech and subtitles, term deletion or substitution, and frame rate perturbation simulation samples.
[0013] In one embodiment of the present application, step (1) further includes an asynchronous compensation mechanism: When the time difference between speech and subtitle exceeds 0.8 seconds, secondary semantic matching is performed based on Jaccard similarity; If the semantic similarity is greater than 0.6, it is considered to be a valid alignment.
[0014] In one embodiment of the present application, the time alignment in step (1) is verified by three-reset reliability weighting, and the calculation formula is: ; in, Indicates semantic text similarity; Indicates the degree of time synchronization; Indicates the degree of overlap between the visual OCR text box and the sound area in the picture. 、 、 Represent different weighting coefficients respectively.
[0015] In one embodiment of the present application, the structured data outputted in step (4) includes: a mapping relationship between frame index and audio timestamp, a term replacement record and a confidence score, and at least one label in the text reconstruction result after semantic coherence optimization.
[0016] By adopting the above technical solution, by building a unified time mapping and dynamic tolerance mechanism, the problem of time drift between video audio, subtitles and image information is effectively solved; with the help of semantic enhancement processing and terminology classification vocabulary design, the accuracy of terminology recognition in professional fields (such as semiconductors, medicine, etc.) is significantly improved; the sample dynamic classification mechanism can achieve training sample quality control, adapt to different training stages, and improve the model's generalization ability and robustness; the final output structured training data is highly consistent and interpretable, significantly reducing the cost of manual labeling and improving training efficiency and model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present invention will be described in detail below with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a flow chart of the first embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the following specific embodiments are only used to explain the present invention and do not constitute a limitation of the present invention.
[0019] like Figure 1 As shown, in order to achieve the above purpose, the present invention proposes a video training data generation method based on multimodal semantic alignment, comprising the following steps: (1) Perform multimodal time alignment on the audio, image frames and text information in the video to establish a cross-modal temporal mapping relationship; (2) Performing semantic enhancement processing based on the time alignment results to improve the recognition accuracy of professional terms; (3) Dynamically classify training samples according to semantic density and confidence; (4) Output hierarchical structured training data to adapt to different training stages.
[0020] Specifically, the present invention proposes a method for generating video training data based on multimodal semantic alignment. This method combines audio, image frames, and text information to construct high-confidence, scalable video training samples, suitable for a variety of semantically intensive model training scenarios. The following detailed description of the technical solutions of the present invention is combined with specific embodiments to enable those skilled in the art to clearly understand and implement the present invention accordingly.
[0021] In this embodiment, the input is raw video data and the output is a structured, hierarchical training data set. The method includes the following steps: First, three key modalities are extracted from the original video: audio stream, image frame sequence, and visual / speech recognition text. An automatic speech recognition (ASR) model is used to extract the timestamp-annotated text sequence from the audio. Optical character recognition (OCR) is used to extract subtitles from the image frames, and the corresponding frame index is recorded.
[0022] Then, a unified timeline mapping function is used to map the timestamps in the speech recognition results to the corresponding image frame numbers. The mapping function is as follows: 𝐹= 𝜏∗𝑓𝑝𝑠; Among them, 𝑓𝑝𝑠 represents the video frame rate, and 𝜏 is the speech recognition output timestamp.
[0023] To improve time alignment accuracy, the system sets a tolerance range: the default tolerance for audio information is ±0.5 seconds, which is automatically extended to ±1.2 seconds if the audio signal-to-noise ratio is low; the frame tolerance range for subtitle information is set based on the subtitle density: the tolerance for sparse subtitle areas is ±10 frames, and the tolerance for dense subtitle areas is ±3 frames.
[0024] If the detected deviation between the audio text and the subtitle text corresponding to the image frame exceeds 0.8 seconds, the asynchronous compensation mechanism is triggered. This mechanism calculates the Jaccard similarity between the two text segments. When the similarity exceeds 0.6, it is considered a valid match and the time point after the secondary alignment is used as the final mapping result.
[0025] In addition, the system integrates semantic text similarity, temporal synchronization, and spatial overlap between the vocalization area and the subtitles to build a three-way confidence scoring mechanism. The matching score calculation formula is as follows: ; in, Indicates semantic text similarity; Indicates the degree of time synchronization; Indicates the degree of overlap between the visual OCR text box and the sound area in the picture. 、 、 Represent different weighting coefficients respectively. Among them, 、 、 The default values are 0.5, 0.3, and 0.2. Semantic similarity is calculated based on the Levenshtein ratio, temporal synchronization is the normalized time difference function, and visual coincidence is calculated based on the overlap ratio between the OCR text box and the vocal area (such as the mouth) in the image.
[0026] After completing the time alignment, the system enters the semantic enhancement processing stage. First, it loads the preset multi-level terminology database, which is divided into the basic terminology layer (L1) and the subdivided domain terminology layer (L2).
[0027] The basic terminology layer is applied by default to all science and technology or humanities videos, such as the terms "nano", "lithography", etc. When the frequency of a term exceeds a preset threshold (for example, 3 times) during the recognition process, or when the context contains specific keywords, the system automatically activates the subdivided term layer (such as "EUV", "FinFET", etc.).
[0028] Term enhancement methods include: exact replacement, fuzzy semantic correction, and semantic disambiguation.
[0029] Example of exact replacement: If a standard term abbreviation or alias (such as "EUV") is identified in the OCR or ASR results, it will be replaced with the standard term "Extreme Ultraviolet Lithography"; An example of fuzzy semantic correction is to recognize “5 nanometers” as “5 nanometers”; An example of semantic disambiguation: For example, the term "GPU" may mean "graphics processor" or "graphics card". The system extracts the word embedding vector through the context window, calculates the cosine similarity between the term and the context, and selects the most relevant semantic output.
[0030] After alignment and semantic enhancement, the system dynamically grades the samples based on their semantic density (such as the number of terms, semantic complexity) and confidence scores. The grading criteria are as follows:
[0031] Finally, the system converts the graded training samples into a structured data format that includes image frame indexes, speech timestamps, recognized text, term labels, alignment confidence, and sample-level labels. This structured data can be directly fed into the subsequent model training pipeline, supporting a phased curriculum training strategy.
[0032] By adopting the above technical solution, by building a unified time mapping and dynamic tolerance mechanism, the problem of time drift between video audio, subtitles and image information is effectively solved; with the help of semantic enhancement processing and terminology classification vocabulary design, the accuracy of terminology recognition in professional fields (such as semiconductors, medicine, etc.) is significantly improved; the sample dynamic classification mechanism can achieve training sample quality control, adapt to different training stages, and improve the model's generalization ability and robustness; the final output structured training data is highly consistent and interpretable, significantly reducing the cost of manual labeling and improving training efficiency and model performance.
[0033] In one embodiment of the present application, the multimodal time alignment in step (1) includes dynamic time window matching, and the tolerance range of the dynamic time window is adaptively adjusted according to at least one of the signal-to-noise ratio of the audio channel, the density of the visual text, and the time offset between the speech and the subtitles.
[0034] Specifically, first, the audio stream, image frame sequence, and text information obtained through speech recognition (ASR) and optical character recognition (OCR) are synchronously extracted from the original video. The system constructs a unified timeline and maps the timestamp τ in the speech recognition result to the corresponding frame index F through a time mapping function: 𝐹= 𝜏∗𝑓𝑝𝑠; To improve cross-modal alignment accuracy, the system introduces a dynamic time window matching mechanism based on this. The tolerance range of the alignment window is adjusted to adapt to the modal differences in different scenarios. The tolerance range of the dynamic time window is adaptively set based on at least one of the following factors: Audio channel signal-to-noise ratio (SNR): When the audio signal-to-noise ratio is high (for example, SNR ≥ 20 dB), the tolerance window can remain within a small range (for example, ±0.5 seconds). When the audio signal-to-noise ratio is low (for example, in the presence of background noise or unclear speech), the tolerance window automatically expands to a larger range (for example, ±1.2 seconds) to mitigate timing inaccuracies caused by ASR recognition deviations.
[0035] Visual text density: In areas with sparse subtitles, there is less text in the visual frame, and the time window tolerance is set to ±10 frames; in areas with dense subtitles, the tolerance window is reduced to ±3 frames to prevent semantic inconsistency caused by excessive cross-frames.
[0036] Time offset between speech and subtitles: The system calculates the offset Δt between the ASR timestamp and the time when the OCR subtitles appear. When the offset |Δt| exceeds 0.8 seconds, the system triggers the "semantic rematching process" and performs Jaccard similarity analysis on the text segments between the modalities. If the similarity is greater than 0.6, the match is considered valid and the corresponding time windows are realigned.
[0037] Employing the aforementioned technical solution, the present invention significantly improves the accuracy and robustness of multimodal data alignment through a dynamic time window matching mechanism. This mechanism adaptively adjusts the tolerance window based on actual video characteristics such as audio signal-to-noise ratio, subtitle density, and inter-modal temporal offset. This effectively addresses the failure of traditional fixed-window methods in scenarios such as frame skipping, frozen frames, and subtitle lag, ensuring alignment flexibility and semantic consistency.
[0038] In one embodiment of the present application, the dynamic time window matching calculates the mapping relationship between audio timestamps and frame indices using the following formula: 𝐹= 𝜏∗𝑓𝑝𝑠; Among them, 𝑓𝑝𝑠 represents the video frame rate, and 𝜏 is the speech recognition output timestamp.
[0039] In this specific embodiment, a key step in dynamic time window matching is achieving accurate mapping between audio time information and image frame indices. To this end, the system introduces a mathematical mapping formula to convert the timestamp τ (in seconds) output by the automatic speech recognition (ASR) module into the corresponding video frame index F (a non-negative integer). This formula accurately locates speech events in the time domain onto image frames in the spatial domain, establishing a foundation for synchronization between the audio and visual modalities. This provides an accurate temporal reference for subsequent cross-modal alignment, semantic enhancement, and training sample generation.
[0040] In the specific implementation process, the system first reads the frame rate (fps) of the input video, and then combines it with the speech timestamp τ output by the ASR to calculate the corresponding frame index F. This frame index is used to select the corresponding frame image and its OCR recognition result in the visual modality to form the basic unit of cross-modal temporal matching.
[0041] This technical solution, by setting a clear frame index calculation formula, enables precise mapping between audio timestamps and image frames, providing fundamental support for dynamic time window matching. This method boasts a simple structure, high computational efficiency, and high reproducibility, ensuring consistent synchronization between audio and image modalities across different video frame rates, significantly improving the accuracy and stability of multimodal alignment.
[0042] In one embodiment of the present application, the semantic enhancement processing in step (2) includes: Performing term replacement on the recognized text using a hierarchical term base activated at the domain knowledge level; The term ambiguity is resolved based on contextual semantic similarity. The calculation formula is: ; in, is the embedding vector of the term, is the embedding vector of the context window.
[0043] In this embodiment, the semantic enhancement process includes two key steps: one is to replace terms in the recognition text based on the hierarchical term library, and the other is to resolve term ambiguity based on contextual semantic similarity.
[0044] The details of the hierarchical term base and term replacement are as follows: The system pre-builds a structured terminology database, which is hierarchically managed according to the professional depth of domain knowledge, including at least the basic layer (Level 1) and the detailed domain layer (Level 2): The Level 1 terminology library includes common general scientific and technological terms, such as "nano", "photolithography", "algorithm", etc., and is activated by default when the system processes science and technology or social humanities videos.
[0045] The Level 2 terminology library contains terms with industry attributes or high professionalism, such as "FinFET", "EUV lithography", "RISC-V architecture", etc. The system automatically activates this level of the vocabulary library when it detects that the corresponding word frequency reaches the set threshold (such as ≥3 times) or the context contains keywords related to the term.
[0046] The system searches for abbreviations, aliases, or ambiguous expressions in automatic speech recognition (ASR) and optical character recognition (OCR) results. Once identified, it replaces the term with the standard term. For example, if the system recognizes "EUV" in text, it replaces it with the full term "extreme ultraviolet lithography."
[0047] The specific content of contextual semantic resolution of term ambiguity is: The system first extracts a context window around the target term's occurrence in the sentence, for example, a window spanning three words or syntactic units before and after it. It then uses a pre-trained language model (such as BERT) to obtain the following two vector representations: and .
[0048] Based on these two vectors, the cosine similarity is calculated to evaluate the semantic tendency of the term in the current context. The calculation formula is: ; The system compares the similarity scores of multiple candidate interpretations, selects the interpretation with the highest score as the standard interpretation in the current context, and outputs it for structured annotation.
[0049] This technical solution significantly improves the recognition accuracy and clarity of specialized terminology in video scenarios by using a hierarchical terminology database for semantic enhancement and contextual semantic similarity calculation to resolve terminology ambiguity. The on-demand activation mechanism for the terminology database balances processing efficiency with professionalism, avoiding unnecessary terminology loading. The semantic disambiguation mechanism leverages contextual information for dynamic discrimination, effectively reducing misidentification of polysemous terms and improving the accuracy of term labels and the reliability of structured information extraction.
[0050] In one embodiment of the present application, the activation conditions of the hierarchical terminology library include: L1-level terms are loaded by default in specific video categories; L2-level terms are activated when the word frequency exceeds a threshold or the context contains related keywords.
[0051] Specifically, in this embodiment, in order to improve the efficiency and accuracy of semantic enhancement processing, the system adopts a terminology library divided by knowledge level and sets differentiated activation conditions to achieve dynamic loading and on-demand use of the terminology library.
[0052] The terminology database is graded according to the generality and professionalism of the terms, and includes at least the following two levels: L1 and L2.
[0053] Level 1 terminology includes basic terms widely applicable to science and technology or social and humanities videos, such as "nano," "algorithm," "resolution," and "energy consumption." These terms are commonly found in video scenes with high semantic generalization and versatility, so the system loads the L1 terminology library by default for certain video categories.
[0054] Specifically, the system automatically determines which categories fall within the enabled range based on the input video's category label, video title keywords, or video type information extracted from metadata (e.g., "science," "education," "documentary," etc.). If these conditions are met, the L1 terminology database is activated and participates in the subsequent term replacement and semantic annotation process of the recognized text.
[0055] Level 2 terminology encompasses specialized and industry-specific terminology, such as "FinFET," "RISC-V," "blood oxygen saturation," and "EUV lithography." To avoid ineffectively loading a large number of specialized terms in general video scenarios, the system sets activation conditions for L2 terminology, including the following two categories: word frequency threshold triggering and contextual keyword association triggering.
[0056] Among them, the word frequency threshold trigger means that when the cumulative frequency of a term or its synonym in the recognized text in a video clip of a certain length exceeds the set threshold (for example, ≥3 times), the system automatically activates the L2-level vocabulary of the field to which the term belongs to improve the recognition, correction and semantic interpretation capabilities of such terms.
[0057] Contextual keyword association triggering means that the system analyzes and identifies the contextual information of the text. If it is found that it contains related keywords in a certain professional field (such as "transistor", "ultraviolet light", "channel mobility", etc.), it can trigger the loading of the corresponding L2-level term subset to support the replacement and ambiguity resolution of terms in this field.
[0058] Through the above mechanism, the system realizes the dynamic scheduling of terminology resources, takes into account the coverage and runtime efficiency of terminology recognition, and effectively supports the automatic structured processing of multi-domain video data.
[0059] This technical solution, by setting clear activation conditions for the terminology library hierarchy, dynamically adjusts the terminology loading strategy based on video category and semantic context, achieving both efficient and professional terminology recognition. The default loading of L1 terms ensures immediate response to common terms, while the on-demand activation mechanism for L2 terms avoids resource redundancy and recognition misuse, significantly improving the system's semantic accuracy and operational performance in multi-domain video processing.
[0060] In one embodiment of the present application, the dynamic classification in step (3) includes: Primary training set: unimodal data with confidence level higher than 0.85; Intermediate training set: cross-modal validation is consistent and the confidence level is between 0.70-0.85; Advanced training set: contains complex spatiotemporal relationships and has a confidence level lower than 0.70, requiring additional discriminant model annotation.
[0061] Specifically, in this implementation, to adapt to the requirements for sample difficulty and semantic depth at different stages of the model training process, the system introduces a dynamic grading mechanism for training samples based on confidence and modal complexity. This mechanism divides the generated training samples into three levels: primary training set, intermediate training set, and advanced training set. The classification criteria for each level are as follows: The primary training set consists of unimodal data, such as only speech recognition text or only image OCR text, with a confidence score higher than 0.85. This type of data features clear audio and video, clear semantics, and a common vocabulary, making it suitable for the initial stages of model training, helping the model establish basic language-vision alignment capabilities.
[0062] The intermediate training set consists of data segments that have been verified to be consistent across modalities. This means that the speech text and image OCR text match temporally and semantically, and are verified as consistent through a multimodal alignment mechanism. The matching confidence level for this data ranges from 0.70 to 0.85. The intermediate training set is suitable for the model's mid-term learning phase, further strengthening modal fusion and terminology understanding capabilities.
[0063] Advanced training sets contain data segments with complex spatiotemporal relationships, such as causal chains of events, nested terminology, screen transitions, and language supplementation. Due to the complex structure and significant asynchrony of these segments, their confidence scores after alignment or semantic processing are below 0.70. To ensure data reliability, these samples require auxiliary annotation using additional discriminant models, such as alignment verification models, terminology verification models, or context consistency models, to confirm the validity of their structured labels.
[0064] The above three training levels are automatically classified by the system according to the confidence threshold and modal information during the data preprocessing stage. The classification labels will be output together with the samples and used for stage-by-stage training management in the subsequent training scheduling controller.
[0065] By dynamically categorizing training samples into elementary, intermediate, and advanced levels, this technical solution accurately matches the model's learning needs at different training stages, enabling progressive control of training sample difficulty. This strategy not only helps improve the model's learning efficiency for low-noise, high-confidence samples, but also provides structured support for the effective utilization of complex samples. In particular, the introduction of a discriminant model-assisted mechanism for advanced training sets further ensures data quality and significantly improves the stability of model training, its generalization capabilities, and its ability to understand complex semantic structures.
[0066] In one embodiment of the present application, step (3) further includes a training phase advancement mechanism, including: When the validation set accuracy rate is less than 1% for three consecutive training cycles, the stage is switched; At the end of the stage, adversarial samples are injected, including at least one of random mismatching of speech and subtitles, term deletion or substitution, and frame rate perturbation simulation samples.
[0067] Specifically, in this embodiment, in order to improve the robustness and generalization ability of the model during training, the system further introduces a training stage advancement mechanism based on the training sample classification to achieve dynamic adjustment of training difficulty and stage transition.
[0068] This mechanism automatically determines whether to switch to a more difficult training phase based on the model's performance on the validation set during training. It also introduces perturbation samples near the end of the phase to improve the model's anti-interference ability. It specifically includes the following two parts: During each training phase, the system continuously monitors changes in the model's accuracy on the validation set. If it detects that the validation set accuracy has increased by less than 1% over three consecutive training cycles, the system automatically determines that the current training phase has converged and triggers a transition to the next training phase. For example, switching from the basic training phase using a primary training set to the modal alignment phase using an intermediate training set, or from the intermediate training phase to the complex semantic fusion phase using an advanced training set. This strategy avoids resource waste and performance stagnation caused by continuous model training on low-difficulty samples.
[0069] At the end of each training phase (e.g., when training approaches the set epoch number or accuracy improvement slows), the system injects adversarial examples into the training data at a set ratio (e.g., 5%) to enhance the model's adaptability to perturbed data. Injected adversarial examples include, but are not limited to, at least one of the following: artificially disrupting the timeline matching relationship between ASR and OCR to simulate semantic drift; randomly deleting keywords from term-dense text or replacing them with semantically similar words to test the model's ability to understand context; and adjusting the original video frame rate to an asynchronous rate (e.g., from 25fps to 18fps or 30fps) to test the model's robustness to time series perturbations.
[0070] The above-mentioned adversarial samples participate in training together with normal samples, and a certain weight is given through the "soft label" mechanism to guide model learning, so that the model can maintain stable output in the face of data defects, alignment errors or semantic ambiguity.
[0071] By implementing this technical solution and introducing a training phase advancement mechanism, the system automatically controls phase transitions based on changes in model performance, preventing overfitting or training stagnation. Furthermore, the injection of diverse adversarial examples at the end of each phase effectively improves the model's tolerance to data perturbations, semantic ambiguity, and timing errors encountered in real-world applications, enhancing the overall robustness and generalization capabilities of the system. This mechanism makes the model training process more adaptable and challenging, significantly improving the final model's practical performance in complex video semantic processing tasks.
[0072] In one embodiment of the present application, step (1) further includes an asynchronous compensation mechanism: When the time difference between speech and subtitle exceeds 0.8 seconds, secondary semantic matching is performed based on Jaccard similarity; If the semantic similarity is greater than 0.6, it is considered to be a valid alignment.
[0073] Specifically, after completing the initial mapping of audio timestamps to image frame indices, the system further detects the time offset between speech and subtitles. This offset is calculated by comparing the timestamps of the automatic speech recognition (ASR) results with the corresponding frame times of the optical character recognition (OCR) results.
[0074] When the time offset between the two is detected to be greater than 0.8 seconds, the system triggers an asynchronous compensation mechanism. This mechanism no longer relies on original time alignment, but instead uses a semantic-based secondary matching method to confirm alignment.
[0075] The specific semantic matching method is as follows: 1. Extract the ASR output text and OCR recognized subtitle text within the time period; 2. Perform Jaccard similarity calculation on the two 3. When the calculated Jaccard similarity result is greater than 0.6, the system determines that the current speech and subtitle text have a high degree of semantic consistency, and then confirms that the alignment is valid and can be output as part of the training sample.
[0076] 4. If the similarity is lower than the threshold, it is judged as a mismatch, and the current data segment will be marked as "needs manual review" or "degraded" to prevent incorrect samples from affecting training accuracy.
[0077] The above technical solution significantly enhances the multimodal alignment system's ability to handle asynchronous data by introducing a secondary semantic matching mechanism based on Jaccard similarity when the deviation between audio and subtitles exceeds a threshold. This solution avoids semantic confusion caused by misalignment while maintaining the semantic accuracy of structured samples. It effectively expands the scope of available data, improves the fault tolerance and practical applicability of training data, and is particularly suitable for processing video footage where subtitles and audio are out of sync.
[0078] In one embodiment of the present application, the time alignment in step (1) is verified by three-reset reliability weighting, and the calculation formula is: ; in, Indicates semantic text similarity; Indicates the degree of time synchronization; Indicates the degree of overlap between the visual OCR text box and the sound area in the picture. 、 、 Represent different weighting coefficients respectively.
[0079] Specifically, the system sets the following three independent scoring indicators: Semantic text similarity, temporal synchronization, and the overlap between the visual OCR text box and the sound area in the picture.
[0080] Semantic text similarity is obtained by comparing the semantic similarity between automatic speech recognition (ASR) text and optical character recognition (OCR) text, and is often calculated using the Levenshtein ratio or word-level Jaccard similarity. This metric reflects the semantic equivalence of cross-modal content.
[0081] Time synchronization refers to the degree of alignment between the timestamp of the ASR recognition result and the frame time of the corresponding OCR subtitle. It uses the normalization formula:
[0082] Among them, Δt is the time offset between speech and subtitles, The maximum tolerance threshold set for the system (e.g. 1.2 seconds). Values closer to 1 indicate a higher degree of synchronization.
[0083] Spatial overlap measures the degree of spatial overlap between the text box identified by OCR in the image and the vocalization area (e.g., the mouth detection box) in the video. This can be quantified by calculating the intersection over union (IoU) of the bounding boxes of the two, or the ratio of the intersection area to the minimum enclosing area.
[0084] After obtaining the above three indicators, the system calculates the total confidence score according to the following weighted formula:
[0085] The system sets a matching threshold (such as 0.7 or 0.75) based on the above confidence score to determine whether the cross-modal alignment is valid and decide whether to output it as a training sample.
[0086] This technical solution significantly improves the accuracy and stability of multimodal data alignment by constructing a three-dimensional confidence-weighted verification mechanism encompassing semantics, temporal sequence, and spatial location. Compared to single-dimensional matching algorithms, this solution can effectively identify misaligned data with semantically consistent but temporally misaligned samples, or temporally consistent but with mismatched articulation regions, thereby improving the structured quality of training data and the robustness of model learning. Furthermore, this scoring mechanism can be used to control the credibility screening of training samples, enabling the automatic extraction of high-quality data.
[0087] In an embodiment of the present application, the structured data output by the step (4) includes at least one of the following labels: mapping relationship between frame index and audio timestamp, term replacement record and confidence score, and text reconstruction result after semantic coherence optimization.
[0088] Specifically, the mapping relationship between frame index and audio timestamp label is used to describe the positional relationship between each timestamp output by automatic speech recognition (ASR) and its corresponding frame in the image frame sequence. The term replacement record and confidence score label records the term standardization replacement behavior occurring in the semantic enhancement process and the confidence score of the segment after the corresponding term replacement. The text reconstruction result after semantic coherence optimization label is used for the content of partial oralization, disorderly sequence or incomplete subtitle fragments. The system can call a pre-trained language model (such as BERT) to perform semantic reconstruction on the recognized text. The system records the text before and after optimization respectively, and labels the optimization degree.
[0089] By outputting the multi-dimensional labels including frame time mapping, term replacement record, confidence score and semantic reconstruction result in the structured training data, the above technical solution not only improves the readability and interpretability of the training samples, but also enables the model to obtain rich auxiliary information in the learning process, thereby enhancing its perception ability of time sequence structure, semantic consistency and term context.
[0090] The above description is only the preferred embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structural transformation made according to the content of the present application specification and drawings, or direct / indirect application in other related technical fields within the inventive concept of the present application is included in the patent protection scope of the present application.
Claims
1. A method for generating video training data based on multimodal semantic alignment, characterized in that: The following steps are involved: (1) Perform multimodal time alignment on the audio, image frames and text information in the video to establish a cross-modal temporal mapping relationship; (2) Performing semantic enhancement processing based on the time alignment results to improve the recognition accuracy of professional terms; (3) Dynamically classify training samples according to semantic density and confidence; (4) Output hierarchical structured training data to adapt to different training stages.
2. The method for generating video training data based on multimodal semantic alignment according to claim 1, wherein: The multimodal time alignment in step (1) includes dynamic time window matching, and the tolerance range of the dynamic time window is adaptively adjusted according to at least one of the signal-to-noise ratio of the audio channel, the density of the visual text, and the time offset between the speech and the subtitle.
3. The method for generating video training data based on multimodal semantic alignment according to claim 2, wherein: The dynamic time window matching calculates the mapping relationship between audio timestamp and frame index using the following formula: 𝐹 = 𝜏∗𝑓𝑝𝑠; Among them, 𝑓𝑝𝑠 represents the video frame rate, and 𝜏 is the speech recognition output timestamp.
4. The method for generating video training data based on multimodal semantic alignment according to claim 1, wherein: The semantic enhancement process in step (2) includes: Performing term replacement on the recognized text using a hierarchical term base activated at the domain knowledge level; The term ambiguity is resolved based on contextual semantic similarity. The calculation formula is: ; in, is the embedding vector of the term, is the embedding vector of the context window.
5. The method for generating video training data based on multimodal semantic alignment according to claim 4, wherein: The activation conditions of the hierarchical term base include: L1-level terms are loaded by default in specific video categories; L2-level terms are activated when the word frequency exceeds a threshold or the context contains related keywords.
6. The method for generating video training data based on multimodal semantic alignment according to claim 1, wherein: The dynamic classification in step (3) includes: Primary training set: unimodal data with confidence level higher than 0.85; Intermediate training set: cross-modal validation is consistent and the confidence level is between 0.70-0.85; Advanced training set: contains complex spatiotemporal relationships and has a confidence level lower than 0.70, requiring additional discriminant model annotation.
7. The method for generating video training data based on multimodal semantic alignment according to claim 1, wherein: Said step (3) also includes a training phase advancement mechanism, including: When the validation set accuracy rate is less than 1% for three consecutive training cycles, the stage is switched; At the end of the stage, adversarial samples are injected, including at least one of random mismatching of speech and subtitles, term deletion or substitution, and frame rate perturbation simulation samples.
8. The method for generating video training data based on multimodal semantic alignment according to claim 1, wherein: The step (1) also includes an asynchronous compensation mechanism: When the time difference between speech and subtitle exceeds 0.8 seconds, secondary semantic matching is performed based on Jaccard similarity; If the semantic similarity is greater than 0.6, it is considered to be a valid alignment.
9. The method for generating video training data based on multimodal semantic alignment according to claim 1, wherein: The time alignment in step (1) is verified by three-reset reliability weighting, and the calculation formula is: ; in, Indicates semantic text similarity; Indicates the degree of time synchronization; Indicates the degree of overlap between the visual OCR text box and the sound area in the picture. 、 、 Represent different weighting coefficients respectively.
10. The method for generating video training data based on multimodal semantic alignment according to claim 1, wherein: The structured data outputted in step (4) includes: a mapping relationship between frame index and audio timestamp, term replacement records and confidence scores, and at least one label in the text reconstruction result after semantic coherence optimization.
Citation Information
Cited By
Multi-modal large model data integration treatment system and method
CN121144855A
Data Integration and Governance System and Method for Multimodal Large Models
CN121144855B
Video content structured processing method and system fusing voice large model and visual semantic large model
CN121166976A
Video stitching and synthesizing method and device, electronic equipment and storage medium
CN121585881A