Advertisement video question and answer method and system based on audio-visual collaborative awareness and chain verification

By employing a method of audiovisual collaborative perception and chain verification, the problem of inaccurate cross-modal audiovisual information fusion in advertising video Q&A was solved, achieving efficient and accurate feature extraction and answer generation, and improving the accuracy and professionalism of the Q&A results.

CN121962369APending Publication Date: 2026-05-01HUNAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN UNIV
Filing Date
2026-04-02
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing advertising video question-answering methods suffer from inaccurate cross-modal audiovisual information fusion, low feature perception efficiency, difficulty in extracting core promotional audiovisual features, inability to accurately respond to fine-grained audiovisual questions, and the generated answers are prone to factual illusions and logical loopholes.

Method used

The method based on audiovisual collaborative perception and chain verification is adopted. Through multimodal data preprocessing, audiovisual collaborative anchoring, multidimensional advertising knowledge base retrieval and chain reasoning with structured semantic constraints, an initial answer is generated and iteratively optimized through a chain verification mechanism with triggering conditions, and finally an accurate answer is obtained.

Benefits of technology

It achieves efficient and accurate perception and fusion of the audiovisual features of advertising videos, improves the logic of question-and-answer reasoning and the accuracy and professionalism of the answers, and meets the actual needs of the industry for in-depth understanding of advertising videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962369A_ABST
    Figure CN121962369A_ABST
Patent Text Reader

Abstract

The invention provides an advertisement video question and answer method and system based on audio-visual collaborative awareness and chain verification, and relates to the technical field of artificial intelligence, the method comprises the steps of obtaining multi-modal data, the multi-modal data comprising to-be-analyzed advertisement video data and a natural language question input by a user; the method comprises the following steps: preprocessing multi-modal data, extracting a voice text in a video, and performing oversampling and semantic-based key frame optimization on the video to obtain a key frame sequence and an aligned voice text; and based on the obtained key frame sequence and the aligned voice text, constructing an audio-visual and character collaborative awareness and self-adaptive anchoring mechanism, and generating a double-flow visual representation containing a global context and a local fine view angle. According to the method, efficient and accurate perception and fusion of audio-visual features of advertisement videos can be realized, and the logicality of question and answer reasoning and the accuracy and professional degree of answers are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an advertising video question-answering method and system based on audiovisual collaborative perception and chain verification. Background Technology

[0002] With the rapid development of the digital advertising industry, advertising videos have become a core medium for businesses to convey their advertising intentions and achieve commercial communication. Advertising video question answering, as a high-level task of video understanding, can accurately answer users' natural language questions about the content, strategies, and intentions of advertising videos, and has been widely applied in scenarios such as advertising effectiveness analysis, intelligent advertising decision-making, and user advertising consultation. The high information density, complex audiovisual modal relationships, and abstract semantics of advertising videos place stringent demands on the cross-modal perception, professional knowledge integration, and logical reasoning capabilities of question answering systems.

[0003] Existing question-answering methods for advertising videos generally suffer from the following drawbacks: inaccurate cross-modal audiovisual information fusion and low efficiency of feature perception. Because existing methods often employ frame-level rigidity or global average alignment for audiovisual modalities, they fail to consider the common asynchronous phenomenon of audio and video in advertising creation and lack fine-grained visual information perception mechanisms. This results in broken semantic connections across modalities, making it difficult to extract core promotional audiovisual features and accurately respond to fine-grained audiovisual combination questions. Furthermore, inefficient feature perception and modal fusion make it difficult for existing systems to extract effective features with high signal-to-noise ratios, hindering precise integration with professional advertising knowledge. Reasoning lacks reliable audiovisual evidence, easily leading to factual illusions and failing to interpret the deeper promotional intentions and persuasive strategies of advertisements. The generated answers are either merely superficial visual descriptions or contradict the actual audiovisual content and promotional logic of the advertisement, reducing the accuracy and professionalism of the question-answering results and failing to meet the industry's actual needs for in-depth understanding of advertising videos. Summary of the Invention

[0004] This invention provides a question-answering method and system for advertising videos based on audiovisual collaborative perception and chain verification. It can achieve efficient and accurate perception and fusion of audiovisual features of advertising videos, improve the logic of question-answering reasoning and the accuracy and professionalism of answers, and meet the actual needs of the industry for in-depth understanding of advertising videos.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, an advertising video question-answering method based on audiovisual collaborative perception and chain verification, the method comprising: Acquire multimodal data, which includes advertising video data to be analyzed and natural language questions input by users; preprocess the multimodal data by extracting speech text from the video, and by oversampling the video and semantically-based keyframe optimization to obtain a keyframe sequence and aligned speech text; Based on the obtained keyframe sequence and aligned speech text, a co-perception and adaptive anchoring mechanism for audiovisual and textual data is constructed to generate a dual-stream visual representation that includes global context and local fine-grained perspective. Based on the natural language questions obtained from user input and the generated dual-stream visual representation, semantic keywords are extracted, and at the same time, a search is performed in a pre-built multi-dimensional advertising knowledge base to obtain expert context information that matches the current context. The generated dual-stream visual representation, the retrieved expert context information, and the user-input natural language question are concatenated into a multimodal input sequence to drive the large language model to perform chained reasoning and answer generation under structured semantic constraints, thereby obtaining the initial answer. Based on the initial answer, a chain-like verification mechanism based on trigger conditions is introduced. Through feedback loop, the initial answer is fact-checked and iteratively optimized to obtain the final answer.

[0006] Furthermore, multimodal data is acquired, including the advertising video data to be analyzed and the natural language questions input by the user; the multimodal data is preprocessed by extracting speech text from the video, and by oversampling the video and performing semantic-based keyframe optimization to obtain a keyframe sequence and aligned speech text, including: Based on the acquired advertising video data to be analyzed, the audio text in the video is extracted. The extracted audio text is then timestamped and cleaned to obtain the aligned audio text. Based on the acquired advertising video data to be analyzed, the video is subjected to high-frequency uniform oversampling to obtain a sampled frame sequence. Pre-trained visual and language models are used to extract image features, semantic vectors of user-input natural language questions, and semantic vectors of aligned speech text from the sampled frame sequence, respectively. At the same time, the cross-modal similarity between the image features of each sampled frame and the semantic vectors of user-input natural language questions and aligned speech text is calculated. Based on the cross-modal similarity, key frames that meet a preset number are selected from the sampled frame sequence to obtain a key frame sequence.

[0007] Furthermore, generating the dual-stream visual representation includes: Based on the keyframe sequence, the visual space feature map of the current keyframe is extracted; at the same time, the semantic vector of the natural language question input by the user, the semantic vector of the aligned speech text, and the semantic vector of the preset OCR keywords are extracted respectively. Through cross-modal interaction, a user question heatmap, a speech heatmap, and an OCR text heatmap are generated. Based on the speech heatmap, to address the issue of audio-visual asynchrony, a time-series sliding window is set with the current keyframe as the center. The attention weights of the speech text to the current visual frame at each moment within the window are aggregated to generate a time-aligned speech heatmap. The user question heatmap and OCR text heatmap are then weighted and fused with the time-aligned speech heatmap to generate a comprehensive audiovisual and text response map. Calculate the spatial weighted average of the audiovisual and text integrated response map to locate the semantic centroid coordinates; with the centroid coordinates as the center, expand outwards in all directions in combination with the original image size and the preset multi-scale set to generate multiple candidate anchor boxes, and truncate the boundary coordinates of each candidate anchor box at the same time. Calculate the normalized response density score within each candidate anchor frame, select the anchor frame with the highest score as the final local visual region for cropping, and obtain a local fine-view image; at the same time, retain the original keyframe image after uniform scaling as the global view image; stitch the global view image and the local fine-view image together to form a two-stream visual representation that includes global context and local fine-view.

[0008] Furthermore, based on the natural language questions input by the user and the generated dual-stream visual representation, semantic keywords are extracted. Simultaneously, a search is performed in a pre-built multi-dimensional advertising knowledge base to obtain expert contextual information matching the current context, including: Based on the natural language questions input by the user and the dual-stream visual representation that includes global context and local fine-grained perspective, semantic keywords are extracted and mapped into text-based search query tags; Based on the search query tags, a search is conducted in a pre-constructed multidimensional advertising knowledge base to obtain multiple candidate knowledge entries that match the current context; wherein, the multidimensional advertising knowledge base includes at least the knowledge dimensions of advertising persuasion strategies, emotional stimulation mechanisms, visual metaphor mapping relationships, and target audience profile characteristics; The multiple candidate knowledge items are reordered by confidence level to select a preset number of knowledge items with the highest confidence level as expert context information.

[0009] Furthermore, the generated dual-stream visual representation, retrieved expert contextual information, and user-input natural language questions are concatenated into a multimodal input sequence to drive the large language model to perform chained reasoning and answer generation under structured semantic constraints, obtaining the initial answer, including: The generated dual-stream visual representation, the filtered expert context information, and the user-input natural language question are concatenated into a multimodal input sequence according to a preset format. A structured semantic constraint decoding mechanism is constructed, and a chain generation protocol of observation, analysis, reasoning, and answer is set to guide the large language model to reason and generate according to the specified cognitive path; Under the structured semantic constraint decoding mechanism, the multimodal input sequence is input into a large language model, and fine-grained semantic constraints are applied during the text generation process of the large language model. The fine-grained semantic constraints include at least the following: action refinement constraints, which are used to map general action words to specific dynamic behavior words; sentiment valence quantification constraints, which are used to extract keywords of emotional atmosphere of the scene; and visual text anchoring constraints, which are used to preferentially use OCR text information visible in the local fine-view image in the dual-stream visual representation. Based on the applied fine-grained semantic constraints, the large language model is driven to perform step-by-step reasoning according to the chain generation protocol of observation, analysis, reasoning, and answer, generating an initial answer that includes the intermediate reasoning process.

[0010] Furthermore, based on the initial answer, a chain-like verification mechanism based on trigger conditions is introduced. Through a feedback loop, the initial answer is fact-checked and iteratively optimized to obtain the final answer, including: Based on the generated initial answer, the system detects whether the user-input natural language question contains preset high-risk triggering keywords. If not, the initial answer is directly output as the final answer. If it is, the verification process is initiated. High-risk triggering keywords include at least keywords pointing to text details, keywords pointing to logical causality, and keywords pointing to speech content. While keeping the parameters of the large language model frozen, fact-checking prompts are constructed, and the user-input natural language question, the generated initial answer, and the generated dual-stream visual representation are input back into the large language model to guide the large language model to backtrack to multimodal evidence including the keyframe sequence and the aligned speech text, in order to evaluate whether there are factual illusions or logical loopholes in the initial answer, and generate feedback text containing confirmation tokens or correction tokens. The generated feedback text is parsed. If the feedback text contains a confirmation token, the initial answer is determined to be correct and is retained as the final answer. If the feedback text contains a correction token, the corrected answer content is extracted from the feedback text as the final answer, thus completing the iterative optimization of the initial answer.

[0011] Secondly, an advertising video question-answering system based on audiovisual collaborative perception and chain verification includes: The data preprocessing module is used to acquire multimodal data, which includes advertising video data to be analyzed and natural language questions input by users; the multimodal data is preprocessed by extracting speech text from the video, and by oversampling the video and semantically-based keyframe optimization to obtain keyframe sequences and aligned speech text; The audiovisual collaborative anchoring module is used to construct an audiovisual and text collaborative perception and adaptive anchoring mechanism based on the obtained keyframe sequence and aligned speech text, and generate a dual-stream visual representation that includes global context and local fine perspective. The knowledge base retrieval module is used to extract semantic keywords based on the natural language questions input by the user and the generated dual-stream visual representations. At the same time, it searches in a pre-built multi-dimensional advertising knowledge base to obtain expert context information that matches the current context. The chain reasoning generation module is used to concatenate the generated dual-stream visual representation, the retrieved expert context information, and the user-input natural language question into a multimodal input sequence, so as to drive the large language model to perform chain reasoning and answer generation under structured semantic constraints and obtain the initial answer; The chain-verification optimization module is used to introduce a chain-verification mechanism based on trigger conditions, based on the initial answer. Through feedback loop, the initial answer is fact-checked and iteratively optimized to obtain the final answer.

[0012] The above-described solution of the present invention has at least the following beneficial effects: Because it employs multimodal data preprocessing techniques to extract aligned speech text and optimize semantic keyframes, it overcomes the problems of messy raw data and difficulty in selecting effective features in advertising videos, thus providing standardized, high signal-to-noise ratio foundational data for feature processing. Because it uses audiovisual collaborative anchoring techniques, it solves the asynchronous audio-visual problem through a time-series sliding window and adaptively generates dual-stream visual representations, overcoming the core problems of inaccurate cross-modal audiovisual information fusion and low efficiency in fine-grained feature perception, thus achieving efficient and accurate extraction and fusion of core audiovisual features of advertising videos. Because it uses knowledge base retrieval techniques to extract semantic keywords and match them with context from high-confidence advertising experts, it overcomes the lack of domain expertise and professional support in reasoning found in general models. The system addresses the issue of supporting visual features and professional knowledge, thus supplementing the question-and-answer reasoning with domain-specific information. By employing a chain-like reasoning method with structured semantic constraints, it generates answers according to a fixed protocol and applies fine-grained semantic constraints, overcoming the problems of unstandardized reasoning processes and easily divergent answer generation. This results in more rigorous reasoning logic and initial answers that better align with the actual content of the advertisement. Furthermore, by using a trigger-based chain-like verification optimization method, it guides the model to backtrack evidence to complete answer verification and correction, overcoming the problems of question-and-answer results easily generating factual illusions, logical loopholes, and lacking self-correction capabilities. This effectively corrects answer errors, significantly improving the accuracy and professionalism of the question-and-answer results, ultimately meeting the industry's practical needs for a deep understanding of advertising videos. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating the advertising video question-answering method based on audiovisual collaborative perception and chain verification provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of an advertising video question-and-answer system based on audiovisual collaborative perception and chain verification provided in an embodiment of the present invention; Figure 3 This is another flowchart illustrating the advertising video question-answering method based on audiovisual collaborative perception and chain verification provided in the embodiments of the present invention; Figure 4 This is a schematic diagram of the structure of the audiovisual collaborative perception and adaptive anchoring module of the advertising video question answering method based on audiovisual collaborative perception and chain verification provided in the embodiments of the present invention; Figure 5 This is a logical diagram of the knowledge-enhanced chain-reflective reasoning module of the advertising video question-answering method based on audiovisual collaborative perception and chain verification provided in the embodiments of the present invention; Figure 6 This is a schematic diagram showing the rigorous accuracy comparison of different models of the advertising video question answering method based on audiovisual collaborative perception and chain verification provided in the embodiments of the present invention on the AdsQA dataset. Detailed Implementation

[0014] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0015] To clearly define the hierarchical correspondence between the macroscopic implementation steps and microscopic implementation sub-steps of the technical solution of this invention, the five core macroscopic implementation steps disclosed in this invention are now precisely corresponding to the detailed S1 to S5 microscopic implementation steps in this specific embodiment: Step 1 (multimodal data acquisition and preprocessing), Step 2 (audiovisual collaborative anchoring to generate dual-stream visual representation), Step 3 (multidimensional advertising knowledge base retrieval), Step 4 (chain reasoning under structured semantic constraints to generate initial answer), and Step 5 (trigger-based chain verification optimization to obtain final answer) described in this invention constitute the core macroscopic implementation process of the advertising video question-answering method of this invention; S1 (multimodal data acquisition and preprocessing), S2 (audiovisual-text collaborative perception and...) described in this specific embodiment... S1 (adaptive anchoring), S2 (advertising and marketing knowledge base retrieval), S3 (chain-based reasoning generation based on structured semantic constraint decoding), and S5 (chain-based verification and optimization based on feedback loop) are detailed implementations of the above five macro steps. They are in a one-to-one hierarchical relationship, that is, S1 corresponds to step 1 of the invention content, S2 corresponds to step 2 of the invention content, S3 corresponds to step 3 of the invention content, S4 corresponds to step 4 of the invention content, and S5 corresponds to step 5 of the invention content. The sub-steps that further break down S1 to S5 in this specific embodiment (such as S11 / S12 / S13, S21-S24, etc.) are all specific implementation means under the corresponding macro steps. The overall technical process is completely consistent with the core method recorded in the invention content, and no technical features have been added, removed, or changed.

[0016] like Figure 1 As shown, embodiments of the present invention propose an advertising video question-answering method based on audiovisual collaborative perception and chain verification, the method comprising the following steps: Step 1: Obtain multimodal data, which includes advertising video data to be analyzed and natural language questions input by the user; preprocess the multimodal data by extracting speech text from the video, and by oversampling the video and semantically-based keyframe optimization to obtain keyframe sequences and aligned speech text. Step 2: Based on the obtained keyframe sequence and aligned speech text, construct a visual and text collaborative perception and adaptive anchoring mechanism to generate a dual-stream visual representation that includes global context and local fine-grained perspective. Step 3: Based on the natural language questions input by the user and the generated dual-stream visual representation, extract semantic keywords and simultaneously search in the pre-built multi-dimensional advertising knowledge base to obtain expert context information that matches the current context. Step 4: The generated dual-stream visual representation, the retrieved expert context information, and the user-input natural language question are concatenated into a multimodal input sequence to drive the large language model to perform chained reasoning and answer generation under structured semantic constraints, thereby obtaining the initial answer; Step 5: Based on the initial answer, a chain-like verification mechanism based on trigger conditions is introduced. The initial answer is fact-checked and iteratively optimized through a feedback loop to obtain the final answer.

[0017] In this embodiment of the invention, multimodal data preprocessing enables the initial screening of effective features of advertising videos, laying a standardized data foundation for subsequent processing steps. By leveraging a collaborative perception and adaptive anchoring mechanism of audiovisual and textual information, cross-modal audiovisual information is accurately fused, efficiently extracting the core visual features of advertising videos and solving the challenges of asynchronous audio-visual interaction and fine-grained feature perception. Combined with retrieval from a multidimensional advertising knowledge base, professional domain knowledge support is provided for question-answering reasoning, compensating for the shortcomings of general models in terms of professional cognition. Through chain-like reasoning under structured semantic constraints, the logic of answer generation is made more rigorous, ensuring that the initial answer better aligns with the actual content and promotional intent of the advertising video. Further chain-like verification and feedback optimization based on trigger conditions effectively avoids factual illusions and logical loopholes, achieving precise correction of the initial answer and reasonably improving the accuracy and professionalism of the advertising video question-answering results. This effectively meets the industry's actual needs for in-depth understanding and accurate question-answering of advertising videos.

[0018] In a preferred embodiment of the present invention, step 1 above may include: Step 1.1: Based on the acquired advertising video data to be analyzed, extract the speech text from the video. The extracted speech text is then timestamped and cleaned to obtain aligned speech text. Specifically, this includes: separating the audio and video tracks of the advertising video data to be analyzed, extracting the audio track data separately, and performing preprocessing on the extracted audio track data, sequentially performing audio noise reduction, volume normalization, and silent segment removal to eliminate background noise, volume fluctuations, and interference from silent segments on subsequent recognition, thereby improving the purity of the audio signal; simultaneously, the preprocessed audio content is transcribed sentence by sentence, converting the continuous audio signal into editable text form. The audio text content is transcribed and its temporal information is preserved. The transcribed audio text is precisely timestamped segment by segment and sentence by sentence. Each content segment of the audio text is matched with the start and end times of the corresponding video timeline, so that the audio text and the playback progress of the advertising video are accurately matched. After the timestamping is completed, the audio text undergoes multi-dimensional cleaning processing. First, meaningless interjections and repetitive expressions are removed from the text. Then, text errors and sentence segmentation deviations that occurred during automatic recognition are manually corrected according to rules. At the same time, the expression format and punctuation are standardized. Finally, the audio text is aligned with the advertising video timeline, with accurate and standardized content and complete temporal information.

[0019] Step 1.2: Based on the acquired advertising video data to be analyzed, the video is subjected to high-frequency uniform oversampling to obtain a sampled frame sequence. Pre-trained visual and language models are used to extract image features, semantic vectors of the user-input natural language questions, and semantic vectors of the aligned speech text from the sampled frame sequence. Simultaneously, the cross-modal similarity between the image features of each sampled frame and the semantic vectors of the user-input natural language questions and the aligned speech text is calculated. Based on the cross-modal similarity, keyframes meeting a preset number are selected from the sampled frame sequence to obtain a keyframe sequence. Specifically, this includes: performing uniform oversampling of the advertising video data to be analyzed throughout the entire time period at a preset high-frequency sampling frequency, where the preset high-frequency sampling frequency is [number of samples per second]. Multiple frames are sampled at a frequency higher than the original frame rate of the advertising video to ensure complete capture of fleeting text information, logos, and other fine-grained key visual information. This breaks the continuous temporal sequence of the video, breaking it into a discretely distributed sequence of sampled frames that fully covers the entire video content. The pre-trained visual and language models used in this step are an improvement on the CLIP visual-language pre-trained model architecture. The core advantage of this architecture is its dual-encoder independent design, which can achieve accurate cross-modal semantic alignment of visual and text features without additional fusion layers. It has strong generalization ability and is suitable for the core needs of audiovisual semantic fusion and cross-modal feature matching in advertising video question answering. After domain-specific improvements, it can quickly fit the exclusive semantic context of advertising.

[0020] This pre-trained visual and language model is built around the question-answering needs of advertising videos. Its core consists of two main modules: a visual encoder and a text encoder. A feature mapping layer serves as the connecting component between the two modules. The functions and relationships of each module are closely aligned with the cross-modal processing scenarios of advertising. The visual encoder is an improvement on the ViT-Large Visual Transformer architecture. Its core function is to extract features layer by layer from advertising video frames. A newly added multi-scale feature fusion layer enhances the ability to capture small-sized visual targets and local details. After adjusting the feature extraction window, it can accurately identify key visual information such as logos, product parameters, and text information in the advertising image, outputting a high-dimensional graph. The text encoder, built on an optimized Transformer architecture, performs semantic encoding on textual data. By incorporating semantic segmentation strategies specific to the advertising field and adding professional vocabulary embedding dimensions, it accurately analyzes users' natural language questions and the core semantics and core appeals in advertising voice text, outputting standardized text semantic vectors. The feature mapping layer, serving as a connecting component between the two modules, maps the image feature vectors output by the visual encoder and the semantic vectors output by the text encoder to the same common semantic space, eliminating feature heterogeneity between modalities, providing a unified feature foundation for cross-modal similarity calculation, and enabling feature linkage between the two encoders.

[0021] The model training phase utilizes a dedicated image-text dataset specific to the advertising video domain as its core training data. This dataset encompasses labeled data such as keyframes and corresponding text from advertising videos across various industries, advertising Q&A corpora, and product description text. General image-text data is also incorporated for auxiliary training. A cross-modal contrastive learning training approach is employed, allowing the visual encoder and text encoder to collaboratively learn the precise matching relationship between advertising visual features and advertising text semantics through a feature mapping layer. During training, labeled data from real-world advertising video Q&A scenarios, such as product information queries, promotional strategy analysis, and advertising intent interpretation, are continuously introduced. The parameters of the three modules are iteratively optimized to reduce the model's error in cross-modal feature matching within the advertising domain, until the model can stably and accurately identify the semantic relationship between advertising video visuals and text content, thus completing the model's pre-training.

[0022] After pre-training, the visual encoder in the model is invoked to extract features layer by layer from each frame in the sampled frame sequence, obtaining high-dimensional image features corresponding to each sampled frame. Simultaneously, the text encoder in the model is invoked to perform semantic encoding on the user-input natural language question and the aligned speech text obtained in step 1.1, respectively. After mapping through the feature mapping layer, the semantic vectors of the natural language question and the speech text in the same common semantic space are obtained. Then, the cross-modal similarity of the image features of each sampled frame combined with the semantic vectors of the natural language question and the speech text is calculated. The sampled frames that are highly related to the user question and the advertising speech content are selected based on the similarity. The sampled frame sequence is sorted in descending order according to the cross-modal similarity. From the sorted sampled frames, the top-K sampled frames (the top-K sampled frames selected after similarity sorting, where the K value is set according to the advertising video duration and information density, generally 16 frames are selected to ensure that the key visual information is complete and without redundancy) are selected and integrated into a key frame sequence that can accurately reflect the core content of the advertising video according to the original video time sequence, providing highly relevant visual basic data for audiovisual collaborative perception.

[0023] In a preferred embodiment of the present invention, step 2 above may include: Step 2.1: Based on the keyframe sequence, extract the visual space feature map of the current keyframe; simultaneously, extract the semantic vector of the user-input natural language question, the semantic vector of the aligned speech text, and the semantic vector of the preset OCR keywords. Through cross-modal interaction, generate a user question heatmap, a speech heatmap, and an OCR text heatmap. Specifically, based on the obtained keyframe sequence, extract the visual space feature map V of the current keyframe frame by frame. V is the high-dimensional visual feature distribution matrix of the current keyframe in pixel space. First, a non-linear mapping layer is used to combine the visual features and the semantic vector of the user-input natural language question. Aligned speech-text semantic vectors and preset OCR keyword semantic vectors All features are projected onto the same common semantic space to eliminate feature heterogeneity between different modalities. Then, the cross-modal attention distribution of visual features and the three types of text vectors is calculated separately. , , For learnable modal projection matrices, The feature dimension scaling factor. For vector dot product operation, For the preset OCR keyword set, This represents the semantic vector of the k-th keyword in the preset OCR keyword set. (or All are sets of spatial locations across the entire image, τ is the temperature coefficient, and the semantic vector for the user's question is defined by the formula. The calculated user problem heatmap is used, with the superscript T representing the vector transpose symbol, which is used to adapt the dimension format of the dot product operation.

[0024] For speech-text semantic vectors, using the formula The speech heatmap was calculated. and Representing the visual feature maps in spatial coordinates respectively To determine the semantic response strength to user questions and voice content, a multi-head semantic maximization activation strategy is adopted based on the preset OCR keyword semantic vector, using a formula... It captures the salient features of keywords such as text and logos in an image, suppresses the interference of background noise on key text regions, and finally generates the corresponding OCR text heatmap.

[0025] Step 2.2: Based on the speech heatmap, to address the audio-visual asynchrony, a time-series sliding window is set with the current keyframe as the center. The attention weights of the speech text to the current visual frame at each moment within the window are aggregated to generate a time-aligned speech heatmap. The user question heatmap and OCR text heatmap are then weighted and fused with the time-aligned speech heatmap to generate a comprehensive audiovisual and text response map. Specifically, this includes: addressing the common audio-visual asynchrony in advertising video creation due to the pursuit of dissemination effects, a time-series sliding window with a radius of w is set with the time step t of the current keyframe as the center. , Visual features at the current time step t, Let j be the speech-text vector at time step j. The cross-modal attention calculation function is obtained through the formula. The cross-modal attention scores of the speech-text vector and the current visual features at each time step within the calculation window are accumulated. The attention weights of the speech-text vector to the current visual frame at each time step within the window are aggregated to achieve temporal alignment between speech semantics and visual frames. A temporally aligned speech heatmap is generated, and then a linear weighted fusion strategy is applied. This is an adjustable modal balance coefficient, corresponding to the contribution weights of user questions, speech, and OCR text modalities, respectively, expressed by the formula... Heatmap of user issues ( OCR text heatmap ( ) and time-series aligned speech heatmap ( The system performs weighted fusion, integrating semantic cues from different modalities to ultimately generate an audiovisual-text integrated response map that comprehensively represents the audiovisual-text semantic response. .

[0026] Step 2.3: Calculate the spatially weighted average of the audiovisual and text integrated response map to locate the semantic centroid coordinates; using the centroid coordinates as the center, expand outwards in all directions based on the original image size and a preset multi-scale set to generate multiple candidate anchor boxes. Simultaneously, truncate the boundary coordinates of each candidate anchor box. Specifically, this includes: based on the generated audiovisual and text integrated response map... H , H(i,j) In response to the graph in spatial coordinates (i,j) The semantic response value at the location is obtained through the formula. , Calculate the spatial weighted average values ​​of the values ​​along the horizontal and vertical axes respectively, and use this to accurately locate the semantic centroid coordinates where the semantic response intensity is densest in the image. Then, using the semantic centroid coordinates as the geometric center, and combining the actual width W and height H of the original image, the scaling factor is calculated according to a preset multi-scale scaling factor set s (the set of values ​​is...). It expands outwards. For candidate anchor boxes at a specific scale s, The boundary constraint function is defined by the following formula: ; Multiple candidate anchor boxes covering different receptive field ranges are generated, and a boundary constraint function is introduced. The boundary coordinates of each candidate anchor box are numerically truncated. Specifically, based on the original boundary coordinates of the generated candidate anchor boxes, coordinate values ​​exceeding the physical boundaries of the image are forcibly constrained within the effective range of the image: if the left boundary coordinate of the anchor box is less than 0, it is corrected to 0; if the right boundary coordinate of the anchor box is greater than the image width, it is corrected to the image width; similarly, if the upper boundary coordinate of the anchor box is less than 0, it is corrected to 0; if the lower boundary coordinate of the anchor box is greater than the image height, it is corrected to the image height. This ensures that the boundary coordinates of all candidate anchor boxes fall within the effective pixel space of the image, avoiding out-of-bounds sampling or invalid regions; and forcibly ensures that the generated anchor box coordinates are strictly limited to the physical boundaries of the image. Within the specified range, invalid out-of-bounds sampling should be avoided.

[0027] Step 2.4: Calculate the normalized response density score within each candidate anchor box, select the anchor box with the highest score as the final local visual region for cropping, and obtain a local fine-view image; simultaneously, retain the original keyframe image after uniform scaling as the global view image; stitch the global view image and the local fine-view image together to form a two-stream visual representation containing global context and local fine-view, specifically including: traversing all generated candidate anchor boxes at all scales, For the normalized response density score of candidate anchor boxes at a specific scale s, the molecular The cumulative sum of heatmap energy within the region represents the absolute semantic intensity; the denominator is... This is a scale penalty factor used to offset the energy deviation caused by differences in anchor frame area; it is expressed by the formula... Calculate the normalized response density score within each candidate anchor frame, and then use the formula... Select the anchor box with the highest score The final local visual region is precisely cropped to obtain a fine-grained local view image that can capture tiny text and detailed features in the image, while retaining the original keyframe image that has been uniformly scaled. , This is an image scaling function. The image cropping function uses the original keyframe image as a global viewpoint image that fully preserves scene context information, and is finally calculated using the following formula: ; By stitching together the global view image and the local fine-view image along the feature dimension, a two-stream visual representation that takes into account both global context and local fine details is formed. .

[0028] In a preferred embodiment of the present invention, step 3 above may include: Step 3.1: Based on the user-input natural language question and the generated dual-stream visual representation containing both global context and fine-grained local perspectives, extract semantic keywords and map these semantic keywords into text-based search query tags. Specifically, this includes: based on the user-input natural language question and the generated dual-stream visual representation containing both global context and fine-grained local perspectives, first perform natural language processing on the natural language question, sequentially completing precise word segmentation, part-of-speech tagging, syntactic analysis, and stop word removal operations. Through part-of-speech filtering, core words such as nouns, verbs, and adjectives are identified, accurately extracting the core elements in the question that point to the advertising content, presentation format, promotional strategy, and communication intent. The text semantic keywords are then analyzed in a layered visual feature analysis of the dual-flow visual representation. From a global perspective, macro visual keywords corresponding to the overall scene, composition, color tone, and emotional atmosphere of the advertisement are extracted. From a detailed local perspective, precise micro visual keywords corresponding to explicit text, product details, brand logos, visual symbols, and core subjects in the image are extracted. The extracted text semantic keywords are then fully merged and integrated with the macro and micro visual keywords. The integrated keyword set is then deduplicated and normalized to unify the expression of words, eliminate the differences in expression of synonyms with different forms and near-synonyms, and form a standardized set of core keywords.

[0029] The default format requirements for knowledge base retrieval are that keyword tags must be in the form of structured key-value pairs, including two core default key names: question dimension key and visual feature key. The question dimension key is a default identifier representing the core needs of the user's question, used to associate with text semantic keywords extracted from natural language questions and confirm the question orientation of the retrieval. The visual feature key is a default identifier representing the core visual information of the advertising video, used to associate with visual keywords parsed from dual-stream visual representations and confirm the visual content orientation of the retrieval. All keywords in the key-value pairs must be in the form of single words or fixed phrases, without complex sentences or redundant modifications, and the overall number of keywords must be limited to adapt to the input specifications of knowledge base vector retrieval.

[0030] According to the preset format requirements, the standardized set of core keywords is subjected to structured mapping processing. First, preset key names for question dimensions are matched to the semantic keywords of the text, and preset key names for visual features are matched to the integrated macro and micro visual keywords. Then, the two types of key-value pairs are combined in a fixed order with the question dimension key first and the visual feature key second. Complex expressions that do not meet the format requirements are eliminated, and all keywords are standardized into single words or fixed phrases to ensure that the number of keywords is within the preset range. Finally, a structured text form of retrieval query tags that conforms to the knowledge base retrieval rules and can be directly used for semantic matching is generated to ensure that the tags can be accurately identified and parsed by the knowledge base retrieval system.

[0031] Step 3.2: Based on the search query tags, a search is conducted in the pre-constructed multi-dimensional advertising knowledge base to obtain multiple candidate knowledge entries that match the current context. The multi-dimensional advertising knowledge base includes at least the knowledge dimensions of advertising persuasion strategies, emotional stimulation mechanisms, visual metaphor mapping relationships, and target audience profile characteristics. Specifically, based on the search query tags, a multi-dimensional precise semantic search is performed in the pre-constructed multi-dimensional advertising knowledge base. The semantic encoding model used is an improvement on the BERT-based pre-trained language model architecture. The core advantage of this architecture is its strong contextual semantic understanding capability, which can accurately capture the deep semantic features of the text. Furthermore, the model is lightweight, has a fast inference speed, and is well-suited to the needs of rapid semantic retrieval in the knowledge base. After customized improvements in the advertising domain, it can better capture the semantic relationships of advertising professional texts.

[0032] The specific construction and training process of this pre-trained semantic coding model aligns with the contextual semantic expansion of the advertising video question answering in this invention: On the BERT-base architecture, the model retains its core bidirectional Transformer encoder structure, optimizes the word embedding layer for the text features of the advertising domain, increases the word embedding dimension of professional terms such as advertising persuasion strategies, visual metaphors, and audience profiles, and strengthens the semantic representation ability of professional terms; at the same time, a domain semantic adaptation layer is added after the encoder to improve the semantic coding accuracy of the model for advertising domain text, and finally constructs a customized semantic coding model adapted to the retrieval of advertising professional knowledge base.

[0033] The training phase uses a professional text dataset in the advertising field as the core training data, covering advertising theory literature, advertising case analysis texts, visual metaphor interpretation materials, audience profiling research reports, etc. At the same time, general image-text semantic matching data is added to assist training. A contrastive learning training method is adopted to allow the model to learn the semantic feature mapping rules of professional advertising texts. The model parameters are iteratively optimized in specific application scenarios such as matching advertising copy with publicity theory and matching visual symbol descriptions with metaphor definitions. This enables the model to accurately convert text information in the advertising field into high-dimensional semantic vectors. After the model pre-training is completed, it is used as a dedicated semantic encoding model for knowledge base retrieval. It adopts the same source model as the knowledge base construction to ensure the consistency of semantic vector mapping.

[0034] First, the structured search query tags containing two preset key names are parsed to extract all standardized keywords from the key-value pairs and integrated into the search text. The search text is then input into the pre-trained semantic encoding model. Through the step-by-step processing of the model's word embedding layer, bidirectional Transformer encoder, and domain semantic adaptation layer, text feature extraction and vector mapping are completed, transforming it into a high-dimensional semantic search vector with the same dimension and semantic space as the semantic vector of the knowledge entries in the knowledge base, ensuring the effectiveness and accuracy of vector matching. Next, a full-database semantic vector matching process is initiated within the multi-dimensional advertising and promotion professional knowledge base. Similarity calculations are performed between the retrieval vectors and the semantic vectors of knowledge entries within each of the four core knowledge dimensions of the knowledge base. These dimensions include: advertising persuasion strategies (the various persuasive methods and logical strategies employed by an advertisement to achieve its communication goals, such as scarcity guidance, authoritative endorsement, and social recognition); emotional stimulation mechanisms (the design logic of advertisements to evoke audience emotions through visuals and audio, such as creating feelings of pleasure, warmth, and urgency); visual metaphor mapping relationships (the corresponding association between visual symbols, visual elements, and the underlying meaning in an advertisement, such as the metaphorical expression of color, patterns, and main image); and target audience profile characteristics (the set of characteristics of the core audience group targeted by the advertisement, including demographic attributes, consumption habits, aesthetic preferences, and pain points). Precise matching within each dimension is prioritized, and then the matching results from each dimension are integrated.

[0035] The process begins by setting a preset semantic similarity threshold. This threshold is a pre-defined similarity threshold based on the semantic matching scenario in the advertising domain, ranging from 0 to 1. It is used to define the semantic matching degree between the retrieval vector and the knowledge item vector. All knowledge items with similarity values ​​higher than this threshold are selected, while irrelevant knowledge content with low matching degree is eliminated. Finally, multiple candidate knowledge items that highly match the specific context of the current advertising video Q&A and the core demands of the user's question are obtained. This multi-dimensional advertising professional knowledge base is a pre-constructed standardized and structured professional knowledge system. All knowledge content is stored in a unified text format. Under the four core knowledge dimensions, there are standardized knowledge items covering corresponding professional theoretical definitions, actual advertising application scenarios, typical industry cases, and core applicable conditions. Furthermore, each knowledge item is pre-generated with a matching high-dimensional semantic vector through the same pre-trained semantic encoding model and stored in the vector retrieval library of the knowledge base to meet the needs of rapid semantic matching.

[0036] Step 3.3 involves re-ranking the multiple candidate knowledge items by confidence level to select a predetermined number of knowledge items with the highest confidence ranking as expert context information. Specifically, this includes: performing multi-dimensional confidence re-ranking and refined screening on the acquired candidate knowledge items; using the vector similarity value obtained during semantic retrieval as the core confidence score basis; assigning the highest score percentage to this value according to a predetermined weight; further refining the score by combining the matching degree between the knowledge item and the current advertising video's industry attributes, core content theme, and dissemination format; and supplementing the score by considering the fit between the applicable scenarios of the knowledge item and the actual application scenarios of the current advertising video, using a formula. Complete the weighted calculation, where Score the overall confidence level of candidate knowledge items. The core score is the vector similarity score. For secondary score revision (industry, theme, and presentation matching degree). To supplement the rating (applicability to the scenario). , , Preset weighting coefficients and satisfying , usually set , , The value is based on the weight priority of semantic similarity core matching, industry theme secondary adaptation, and application scenario supplementary verification in the advertising video question answering scenario. It was obtained through a large number of experiments and iterations on the AdsQA dataset and is the optimal weight ratio in this scenario. Finally, the comprehensive confidence score of each candidate knowledge item is obtained to ensure the comprehensiveness and accuracy of the score.

[0037] Next, all candidate knowledge items are arranged in descending order of comprehensive confidence score to form a structured candidate knowledge item ranking list. Each item is labeled with its corresponding knowledge dimension and core matching point for easy reference in subsequent screening. Based on the actual needs of chain reasoning in advertising video Q&A, the context window capacity of the subsequent large model, and the requirement for comprehensive knowledge support dimensions, a preset number of expert context items is set. This ensures that the selected knowledge items provide sufficient professional support for reasoning while avoiding context redundancy due to too many items. From the ranked candidate knowledge item list, the preset number of knowledge items with the highest confidence ranking are selected. These selected knowledge items undergo content structuring integration and simplification, removing redundant descriptions, duplicate cases, and non-core explanations, while retaining professional theoretical definitions, core application points, and key case information matching the current advertisement. Simultaneously, the integrated content is categorized and organized according to its knowledge dimension, forming clear, accurate, and concise expert context information to provide professional and relevant theoretical and case support for chain reasoning in advertising video Q&A.

[0038] In a preferred embodiment of the present invention, step 4 above may include: Step 4.1 involves concatenating the generated dual-stream visual representation, the selected expert context information, and the user-input natural language question into a multimodal input sequence according to a preset format. Specifically, this includes: sequentially concatenating the generated dual-stream visual representation containing global context and detailed local perspectives, the selected expert context information with clear dimensions and concise content, and the user's original input natural language question, according to the input specifications of the multimodal large language model and the reasoning logic of advertising video question answering, using a preset multimodal input format. The preset multimodal input format is a structured sequence format, employing a fixed order of text guidance + professional knowledge + visual features. Each part is labeled with a unique identifier for differentiation. The text part uses a unified character encoding format, and the visual feature part uses model-adapted features. Tensor format ensures that information from each modality can be accurately identified and parsed by the model. First, the user's natural language question is placed at the beginning as the core guiding part of the input sequence, and a question guidance label is added to determine the model's reasoning goal and answer direction. Then, the expert context information, which has been classified and sorted according to knowledge dimensions, is connected to it, and an expert knowledge label is added to provide professional theoretical support for the model. Finally, the size-normalized dual-stream visual representation is concatenated at the end in a visual feature input form that the model can parse, and a visual feature label is added. The visual features of the global perspective and the local fine perspective are kept independent and labeled separately for the global perspective and the local fine perspective, forming a complete multimodal input sequence containing textual information and visual features. This ensures that the large language model can accurately parse each part of the information and establish effective semantic associations.

[0039] Step 4.2: Construct a structured semantic constraint decoding mechanism and set up a chain generation protocol for observation, analysis, reasoning, and answer to guide the large language model to reason and generate according to the specified cognitive path. Specifically, the large language model used in this invention is an improvement on the multimodal large language model Qwen3-VL architecture. The core advantage of this architecture is that it has both efficient visual feature parsing capabilities and text logic reasoning capabilities, supports high-resolution visual input and structured instruction compliance, and has excellent semantic fusion capabilities for multimodal information. After customized improvements in the field of video question answering, it can better adapt to the audiovisual collaborative reasoning and professional knowledge fusion needs of advertising video question answering.

[0040] The specific construction and training process of the large language model is as follows: Based on the Qwen3-VL architecture, the model retains its core structure of visual encoder, text encoder and cross-modal fusion decoder. For the multimodal features of advertising videos, the local detail perception module of the visual encoder is optimized to enhance the feature extraction capabilities of OCR text, product details and logo patterns in the image; at the same time, the cross-modal fusion decoder is upgraded to add a professional knowledge fusion layer in the video question answering domain, so as to achieve efficient semantic fusion of dual-stream visual representation and advertising professional knowledge. A structured generation instruction adaptation layer is also added to the model to improve the model's compliance with the instruction of observation, analysis, reasoning and answer chain generation protocol. Finally, a customized large language model adapted to multimodal question answering of advertising videos is constructed.

[0041] During the training phase, a dedicated dataset for advertising video question answering is used as the core training data. This dataset includes the AdsQA benchmark dataset and audiovisual question answering samples from various industry advertisements. Advertising expertise and multimodal reasoning samples are also incorporated to assist training. A fine-tuning training method is employed to allow the model to learn the analytical patterns of advertising video audiovisual information and the fusion logic of professional knowledge. The model parameters are iteratively optimized in specific application scenarios such as advertising content interpretation, communication strategy reasoning, and visual metaphor analysis. This enables the model to accurately follow structured generation instructions and complete the entire question-and-answer process from audiovisual perception to professional reasoning. After training, this model serves as the core reasoning engine for the advertising video question answering system of this invention.

[0042] The structured semantic constraint decoding mechanism is a standardized decoding framework for advertising video question-answering reasoning built on a customized large language model. This mechanism aims to constrain the model's generation logic and improve the interpretability of reasoning. It replaces the unconstrained, free-generation decoding mode of large language models through three layers of constraint logic: generation rule constraints, stage task constraints, and semantic output constraints. The generation rule constraints ensure that the model must follow a pre-defined cognitive path to complete reasoning, prohibiting skipping steps or divergent generation. The stage task constraints define precise task boundaries and output requirements for each stage of reasoning, prohibiting cross-stage derivation and subjective assumptions. The semantic output constraints limit the semantic precision and expression of the model's generated content, ensuring a high degree of consistency between the content and the input multimodal information and expert context.

[0043] Based on this mechanism, a four-stage chain generation protocol is specifically set up, consisting of observation, analysis, reasoning, and answer. Each stage of the protocol sets precise and specific generation tasks and content output requirements. The observation stage requires the model to objectively extract real factual evidence from the audiovisual level based solely on the input multimodal information, without making any subjective inferences. The analysis stage requires the model to combine expert contextual information to establish a clear logical mapping relationship between the observed factual evidence and professional knowledge in the advertising field. The reasoning stage requires the model to deduce the deep communication strategies, emotional appeals, and expressive intentions behind the advertisement layer by layer based on the established logical connections. The answer stage requires the model to generate conclusive content that accurately matches the user's natural language questions based on the complete reasoning process. This guides the large language model to strictly follow the specified cognitive path to perform step-by-step reasoning and standardized content generation.

[0044] Step 4.3: Under the structured semantic constraint decoding mechanism, the multimodal input sequence is input into the large language model, and fine-grained semantic constraints are applied during the text generation process of the large language model. The fine-grained semantic constraints include at least: action refinement constraints, used to map general action words to specific dynamic behavior words; sentiment valence quantification constraints, used to extract keywords of emotional atmosphere of the scene; and visual text anchoring constraints, used to preferentially use OCR text information visible in the local fine-view image in the dual-stream visual representation. Specifically, under the overall constraint framework of the constructed structured semantic constraint decoding mechanism, the complete multimodal input sequence is completely and accurately input into the customized large language model adapted to multimodal reasoning of advertising videos, ensuring that the text encoding and visual feature format of the input sequence are completely matched with the model input specifications, and successfully triggering the model's text generation and progressive reasoning process.

[0045] Simultaneously, throughout the entire process of text generation in the large language model according to the observation, analysis, reasoning, and answer chain generation protocol, fine-grained semantic constraints are applied comprehensively and specifically to ensure that the constraints are applied throughout every stage of reasoning and generation, avoiding inference bias or content divergence in the model. These fine-grained semantic constraints include at least three specific and implementable constraint requirements. Among them, the action refinement constraint requires the model to abandon vague general action words and accurately map general action words describing advertising images to specific dynamic behavior words that fit the actual performance of the images. This improves the granularity and accuracy of behavior descriptions by combining image details, avoiding vague expressions; sentiment valence... The quantitative constraints require the model to accurately extract the emotional atmosphere keywords conveyed by the color, composition, and subject action of the image from the global perspective of the dual-stream visual representation. Combined with the relevant knowledge of the emotional stimulation mechanism in the expert context information, the model can achieve explicit and standardized expression of the emotion in the image, ensuring that the emotional description fits the actual context of the advertisement. The visual text anchoring constraints require the model to prioritize the use of OCR text information visible in the local fine-view image of the dual-stream visual representation when generating content at each stage. The model should accurately label the OCR text, confirm its source, and use it as conclusive factual evidence for the reasoning process and conclusion output, so as to avoid unfounded subjective inferences as much as possible.

[0046] Step 4.4, based on the applied fine-grained semantic constraints, drives the large language model to perform step-by-step reasoning according to the chain generation protocol of observation, analysis, reasoning, and answer, generating an initial answer that includes intermediate reasoning processes. Specifically, based on the various fine-grained semantic constraints applied to the large language model in all aspects, continuously drives the model to strictly follow the four-stage chain generation protocol of observation, analysis, reasoning, and answer, and carry out step-by-step, progressive, and standardized reasoning to ensure that the reasoning logic of each stage is coherent and the content fits the needs of advertising video Q&A. In the observation stage, the large language model objectively and comprehensively extracts the core visual elements in the advertising image, the OCR text information clearly visible from a local fine perspective, and the key expressions in the audio text based on the input dual-stream visual representation and audio text information. The extracted information is deduplicated and organized to form complete, authentic, and unsubjected factual observation evidence, and the source of each piece of evidence is marked to ensure that the evidence is traceable.

[0047] In the analysis stage, in combination with the screened expert context information, establish an exact logical mapping relationship between each observed factual evidence and the professional knowledge in the advertising field, such as advertising persuasion strategies, visual metaphor mapping relationships, and emotional stimulation mechanisms, analyze the correlation points between the evidence and the professional knowledge, confirm the inference direction that each piece of evidence can support, and avoid the disconnection between the evidence and the knowledge; in the reasoning stage, integrate multi-source audio-visual factual information and professional theoretical knowledge, and in combination with the requirements of fine-grained semantic constraints, gradually deduce the core communication strategy, emotional appeal orientation, and target audience group positioning behind the advertisement layer by layer, clearly present the logical chain of the reasoning, ensure that the reasoning process is interpretable and verifiable, and there is no jumpy reasoning or unfounded derivation; in the answer stage, based on the previous complete and traceable reasoning process, eliminate redundant expressions, focus on the core诉求 of the user's natural language question, generate accurate, concise, and context-appropriate conclusive content, and at the same time retain the key evidence and logical relationships in the reasoning process; finally, according to the set output format, integrate the standardized content output of the four stages, and generate an initial answer that includes a complete intermediate reasoning process, is logically rigorous, evidence-sufficient, and meets the requirements.

[0048] In a preferred embodiment of the present invention, the above step 5 may include: Step 5.1, based on the generated initial answer, detect whether the natural language question input by the user contains a preset high-risk trigger keyword; if not, directly output the initial answer as the final answer; if so, start the verification process; where the high-risk trigger keyword at least includes keywords pointing to text details, keywords pointing to logical causality, and keywords pointing to voice content, specifically including: based on the generated initial answer that includes a complete intermediate reasoning process and is logically rigorous, preprocess the natural language question input by the user, first perform word segmentation operation, and then剔除 the, de, and, and other auxiliary words and conjunctions without actual semantics such as and, or, etc., and only retain the effective words that can reflect the core诉求 of the user; then start the preset high-risk trigger keyword matching detection mechanism, which is a standardized detection logic customized for the advertising video Q&A scenario, and adopts a dual mode of exact match + semantic approximate match. First, exactly match the preprocessed effective words and phrases with the words in the high-risk trigger keyword library to accurately identify the keywords with exactly the same expression, and then through simple semantic comparison, identify the approximate keywords with slightly different expressions but the same core meaning. At the same time, set the matching threshold and verification logic to avoid missed detection and false detection, and ensure the comprehensiveness and accuracy of the matching detection; exactly match the preprocessed effective words with the preconfigured high-risk trigger keyword library customized for the advertising video Q&A scenario word by word and phrase by phrase to ensure that no potential high-risk keywords are missed.

[0049] The high-risk trigger keyword database contains at least three types of core keywords: First, keywords pointing to text details, specifically those asking about details of the OCR text, subtitles, and textual expressions in the ad, such as text content, what the subtitles are, and specific copy; second, keywords pointing to logical causality, specifically those asking about causal relationships and reasoning between ad content, such as why, because of what, and what the basis is; and third, keywords pointing to audio content, specifically those asking about audio text, voice-over content, and audio details in the ad, such as voice-over content, what the voice is saying, and audio details. During matching and detection, if no high-risk trigger keywords are found in the user's question, the initial answer does not require further verification and can be directly output as a compliant, reliable, and user-relevant final answer. If any high-risk trigger keyword is detected in the user's question, a multimodal fact-checking process is automatically initiated to further verify and evaluate the initial answer.

[0050] Step 5.2: While keeping the parameters of the large language model frozen, construct fact-checking prompts. Input the user-input natural language question, the generated initial answer, and the generated dual-stream visual representation back into the large language model to guide it to backtrack to multimodal evidence, including the keyframe sequence and aligned speech text. This is used to evaluate whether there are factual illusions or logical flaws in the initial answer, generating feedback text containing confirmation or correction tokens. Specifically, this involves: keeping all parameters of the customized large language model used in Step 4 frozen, without any parameter updates, fine-tuning, or modifications, ensuring the model's reasoning logic remains consistent and avoiding any impact on the accuracy of fact-checking due to parameter changes; then, based on the types of high-risk keywords triggered in the user question and the specific sources of multimodal evidence, construct prompts specifically for fact-checking. These prompts will precisely guide the model to focus on the key points of verification corresponding to high-risk questions. For example, if the triggered keyword is a text detail, the prompt will guide the model to focus on verifying the initial answer's description of OCR text and subtitles; if the triggered keyword is a logical causal keyword, the prompt will guide the model to focus on verifying the reasoning logic and causal relationships of the initial answer.

[0051] The user's original natural language question, the generated initial answer (including the complete intermediate reasoning process), and the generated dual-stream visual representation containing global context and local fine-grained perspective are all input back into the same large language model to ensure that the input multimodal information is complete and the format is consistent with the model input specifications. The large language model is guided by prompt words to backtrack and align the original multimodal evidence, which includes the advertising keyframe sequence extracted in step 2, the time-aligned speech text, and the on-screen OCR text extracted from the local fine-grained perspective. The model then uses this original evidence to evaluate the initial answer item by item.

[0052] The assessment focuses on three core dimensions to ensure no omissions: First, whether the initial answer contains factual illusions, such as fabricating visual elements, OCR text, audio content, or detailed descriptions not appearing in the advertisement; second, whether the initial answer contains content deviations, such as misreading OCR text, distorting audio content, misidentifying visual elements, or describing the advertisement content inconsistent with the original evidence; and third, whether the initial answer contains logical flaws, such as reversed causal relationships, unfounded reasoning processes, mismatched evidence and conclusions, or excessive extrapolation beyond the scope of the original evidence. After the assessment, a feedback text containing the precise judgment result and assessment basis is generated. The feedback text carries a unique confirmation token or correction token, both of which are preset fixed identifiers used to visually indicate the verification result of the initial answer. The confirmation token indicates that the initial answer has passed the verification, corresponding to a verification result without factual errors or logical flaws, while the correction token indicates that the initial answer needs optimization, corresponding to a verification result with factual errors, content deviations, or logical flaws. The token format is concise and uniform, facilitating rapid parsing and identification in subsequent steps.

[0053] Step 5.3: Parse the generated feedback text. If the feedback text contains a confirmation token, the initial answer is deemed correct and retained as the final answer. If the feedback text contains a correction token, the corrected answer content is extracted from the feedback text and used as the final answer, completing the iterative optimization of the initial answer. Specifically, this includes: parsing the feedback text generated by the large language model sentence by sentence, focusing on identifying the type of judgment token carried in the text, ensuring accurate differentiation between confirmation tokens and correction tokens, and avoiding deviations in subsequent processes due to incorrect token identification. If the parsing reveals that the feedback text contains a confirmation token, it indicates that the initial answer has been verified by multimodal original evidence, the factual description is accurate, the logic is rigorous, and there are no factual illusions or logical loopholes. In this case, the initial answer is directly retained and output as the final answer, ensuring the reliability and compliance of the answer.

[0054] If the feedback text contains a correction token after analysis, the initial answer is determined to contain factual errors, content deviations, or logical flaws, requiring iterative optimization. The corrected answer content, verified by multimodal evidence, is precisely extracted from the feedback text. During extraction, redundant judgment information is removed, retaining only the corrected core statements. This ensures the corrected content closely aligns with the user's core concerns and is completely consistent with the original multimodal evidence. The extracted corrected content then replaces the corresponding erroneous content in the initial answer. After replacement, the corrected answer is briefly reviewed to ensure smooth expression, logical coherence, and natural connection with the error-free reasoning and evidence citations in the initial answer. This process forms the optimized final answer, completing the iterative optimization of the initial answer and further improving its accuracy and reliability.

[0055] The specific implementation process of the embodiments of the present invention also includes the following method steps: like Figure 3 The embodiments of this invention also provide an advertising video question-answering method based on audiovisual collaborative perception and chain verification. This method utilizes the CLIP visual-language pre-trained model for audiovisual collaborative feature perception and adaptive anchoring, combines a marketing knowledge base to construct expert context, and uses the Qwen3-VL multimodal large model to perform chain reasoning and self-reflection verification based on structured semantic constraints. The specific operation process is as follows: Step S1, multimodal data acquisition and preprocessing, aims to use the CLIP model to perform semantic filtering on the original video and extract keyframe sequences with high signal-to-noise ratio, including: S11, Data Loading and Speech Transcription: Receives the advertising video file to be analyzed and the user's input natural language questions. Automatic speech recognition technology is used to extract the speech content from the video, followed by timestamp alignment and noise reduction to generate a speech-to-text sequence.

[0056] S12, Semantic-based keyframe selection: The advertising video is oversampled at high frequency and uniformly. The visual encoder of the CLIP-ViT-Large model is then used to extract image features from the sampled frames, while simultaneously extracting semantic vectors from the user's question and the spoken text. The cosine similarity score between the image features and the question + speech combination vector for each frame is calculated. Based on the scores, the Top-K (K=16) frames are selected as the keyframe sequence and their size is normalized.

[0057] S13, Question Reception: Receive natural language questions from users regarding advertising content, strategies, or intent.

[0058] Step S2, audiovisual-text collaborative perception and adaptive anchoring, achieves adaptive compression of visual features by constructing a three-stream feature extraction mechanism and a dynamic anchor box generation strategy, including: The core implementation process of this step corresponds to Figure 4 (Structural diagram of the audiovisual collaborative perception and adaptive anchoring module), with specific implementation process combined. Figure 4 and the expansion of the given formula: Figure 4 The entire process of dual-stream attention computation, temporal sliding window fusion, and adaptive anchoring and dual-stream visual representation construction based on response density is fully demonstrated. Taking keyframe sequences as input, corresponding heatmaps are generated through visual-question, visual-audio, and visual-OCR three-stream attention computation. After solving the problem of asynchronous audio and video through temporal sliding window, multimodal weighted fusion is completed, and finally, a global keyframe sequence and a local fine sequence are generated as dual-stream features for subsequent inference. The specific implementation steps correspond one-to-one with the technical process and formulas of step S2, which is a visual implementation of step S2.

[0059] S21, Three-stream feature extraction: Extract the visual space feature map V of the current keyframe, project each modality feature onto the common semantic space through a nonlinear mapping layer, and calculate the visual features and the user question vector respectively. ASR speech-text vectors and preset OCR semantic vectors The cross-modal attention distribution, corresponding to Figure 4 The three-stream attention calculation process involves visual-question attention, visual-audio attention, and visual-OCR attention. For the OCR semantic vector, a multi-head semantic maximization activation strategy is employed to capture salient features from keyword sets such as text, logo, and brand. The established formula for calculating the three-stream feature heatmap is as follows: ;

[0060] ; In the above formula, and Quantitatively representing visual feature maps in specific spatial coordinates The semantic response strength at each location to user questions and voice content is determined by analyzing the local visual feature vector at that location. Cross-modal interaction with text or speech vectors; to address heterogeneity between different modalities, a learnable projection matrix is ​​introduced. Features from different modalities are mapped to a unified common semantic space to achieve feature alignment, and similarity is calculated through vector inner product, supplemented by a feature dimension scaling factor. To prevent gradient vanishing due to excessively large dot product values; in OCR feature extraction, utilize a pre-defined keyword set. Set of all spatial locations in the entire map A global scan is performed, and the distribution sharpness of spatial Softmax normalization is adjusted by the temperature coefficient τ to effectively suppress background noise and highlight the salience of key text areas in the image.

[0061] S22, Multimodal Weighted Fusion: A linear weighting strategy is used to fuse the extracted user question heatmap, voice heatmap, and OCR heatmap. Figure 4 In the multimodal weighted fusion stage, to address the audio-visual asynchrony in advertising videos, a temporal sliding window centered on the current keyframe is set up. The attention weights of the audio and text within the window to the current visual frame are aggregated to generate a temporally aligned audio heatmap. Figure 4 The core processing mechanism of the temporal sliding window, based on the established formulas for the aggregation calculation of the audiovisual-text integrated response map and the speech stream temporal sliding window, is as follows: ; ; In the above formula, a linear weighting strategy is used to transform the user problem heatmap. Time-aligned speech heatmap and OCR text heatmap Superimposed and fused semantic cues from different modalities, among which The preset balance coefficient is used to adjust the contribution of each modality to the final result according to task requirements; the temporal sliding window is centered on the current visual frame time step t and has a radius of w, covering a time range. The speech-text vector at each moment within this range is calculated by accumulating the data. With current visual features Cross-modal attention score It captures all relevant speech and semantics of the current scene within the time-axis neighborhood, generates a time-robust speech heatmap, and solves the problem of asynchronous audio and video in advertising videos.

[0062] S23, Multi-scale Dynamic Anchor Box Generation: Based on the generated audiovisual-text integrated response map, its semantic centroid in space is calculated as the core anchor point for visual attention, corresponding to... Figure 4 The preliminary step in generating the dynamic anchor frame; centered on this center of gravity, based on a preset set of scales. Multiple sets of candidate anchor frames are generated, and a boundary truncation mechanism is introduced to ensure that the anchor frames do not exceed the boundaries. The established formulas for centroid positioning and anchor frame generation are as follows: ; In the above formula, the centroid coordinates where the visual semantic response intensity is most concentrated are determined by calculating the weighted average of the comprehensive response map in the horizontal and vertical axes. This serves as the core anchor point for subsequent cropping; then, using this centroid as the geometric center, and combining the width W and height H of the original image, the cropping is performed according to a preset multi-scale scaling factor s (the set of values ​​is usually...). Expand outwards to generate candidate anchor box regions covering different receptive field ranges. ; through boundary constraint functions Numerical truncation is applied to the calculated boundary coordinates to force the anchor frame coordinates to be confined to the physical boundaries of the image. Within the specified range, avoid invalid out-of-bounds sampling.

[0063] S24, Dual-stream representation construction: Traverse all generated candidate anchor boxes, calculate the normalized response density within each anchor box to eliminate energy bias caused by area differences between anchor boxes of different scales, select the anchor box region with the highest response density for cropping as the local fine-grained view, and retain the original image after uniform scaling as the global view. The global keyframe sequence and the local fine-grained sequence generated in this step are... Figure 4The final output, used as the dual-stream feature for subsequent inference, is based on the established formulas for calculating the response density and constructing the dual-stream representation: ; In the above formula, The normalized semantic response density score of candidate anchor boxes at a specific scale s, molecule The cumulative sum of heatmap energy within this region, representing absolute semantic intensity, is the denominator. As a scale penalty factor, it offsets the energy deviation caused by the increase in the anchor frame area, ensuring that small targets can be evaluated fairly; To maximize the optimal local visible area selected through operation, close-up areas containing key text and logos can be accurately captured; the final dual-stream visual representation input. From the original keyframes that have been uniformly scaled It is composed of a (global view) and a local fine-grained image (local view) cropped based on the best anchor frame, which can achieve accurate perception of tiny details while preserving scene context information.

[0064] Step S3, advertising and marketing knowledge base retrieval, including: S31, Knowledge Base Construction: Pre-build a structured JSON knowledge base, whose knowledge dimensions include at least: advertising persuasion strategies (such as scarcity, authoritative endorsement), emotional stimulation mechanisms, visual metaphor mapping table, and target audience profile characteristics.

[0065] S32, RAG retrieval: Extract keywords from user questions and video content, map the refined visual tokens output by S2 into semantic tags in text form, then concatenate these tags with user question keywords, and use retrieval enhancement generation technology to match the most relevant marketing theory entries in the knowledge base.

[0066] S33, Contextual Filtering: Re-rank the search results by confidence level, filter out noisy knowledge, and use high-confidence marketing definitions as expert contextual information.

[0067] Step S4, chain-based reasoning generation based on structured semantic constraint decoding, includes: This step inputs the dual-stream visual features from S2, the expert context from S3, and the user question into the multimodal large language model to perform structured reasoning. This step, together with step S5, constitutes... Figure 5 The core content of (the logical diagram of the knowledge-enhanced chain-based reflective reasoning module) Figure 5 This is a visual representation of the structured reasoning in this step and the chain verification in step S5.

[0068] Figure 5This paper fully demonstrates the complete logical process of marketing knowledge base retrieval, structured main-line reasoning, and self-reflection verification based on multi-dimensional indicators. Taking the user's question as the core input, after obtaining the expert context through marketing knowledge base retrieval, it combines ASR voice, OCR text, and dual-stream visual features to perform structured main-line reasoning of "observation-analysis-reasoning-answer" to generate an initial answer. After self-reflection verification through multi-dimensional indicators, the final optimized answer is output. This logic is completely consistent with the structured reasoning in step S4 and the chain verification in step S5, and is a visual implementation of the core reasoning and verification mechanism of this invention.

[0069] S41, Constructing a Structured Semantic Constraint Decoding Mechanism: Abandoning the freely generated question-and-answer model, constructing a structured generation protocol for observation, analysis, reasoning, and answers, corresponding to... Figure 5 The core path of structured mainline reasoning in China: the guiding model first outputs audiovisual observation evidence, then establishes a logical mapping between the evidence and the question, and finally generates a conclusion.

[0070] S42, applying fine-grained semantic constraints: During the generation process, explicit instructions restrict the model's output to specific semantics. These constraints include at least: action refinement constraints (forcing the mapping of general action words such as "run" to specific behavior words such as "sprint"), sentiment valence quantification constraints (explicitly extracting keywords related to the emotional atmosphere of the scene), and forced visual text anchoring (based on the local visual input of S2, prioritizing the use of visible OCR text information). This constraint mechanism is implemented throughout... Figure 5 The entire process of structured main-line reasoning ensures the accuracy and standardization of reasoning.

[0071] S43, Multi-source Context Joint Reasoning: Combining the dual-stream visual features (global perspective and local fine-grained perspective) generated in S2 and the marketing expert context retrieved in S3, chained reasoning is performed under the above semantic constraints to generate an initial answer containing intermediate reasoning processes, corresponding to... Figure 5 The initial answer generation stage provides a foundation for subsequent verification.

[0072] Step S5, based on feedback loop chain verification and optimization, includes: This step utilizes a trigger-based verification proxy mechanism to perform dynamic error correction during the inference phase. Figure 5 The specific technical implementation of the self-reflection and verification closed loop, together with step S4, completes the entire process of knowledge-enhanced chain-like reflective reasoning.

[0073] S51, Keyword-based trigger condition detection: Detects whether the user's question contains a preset set of high-risk trigger words. This set includes at least keywords pointing to text details (text, logo, number), keywords pointing to logical causality (why, reason), and keywords pointing to the spoken content (say). If present, the current question is determined to be a complex reasoning task, triggering the S52 verification process. Figure 5 The verification triggers the process; otherwise, the initial answer is output directly.

[0074] S52, Closed-loop reflection based on verification proxy: If verification is triggered, keep the model parameters frozen, construct the verification proxy module, and correspondingly... Figure 5 The core role of the fact checker constructed in the model (the core role of the fact checker is: the concrete execution subject of the self-reflection and verification closed loop of the knowledge-enhanced chain-reflective reasoning module, whose functions, actions, and goals are highly consistent with the module logic) involves inputting user questions and the initial answers generated by S4 as drafts to be verified, along with visual features, back into the model to activate the fact-checking logic. This guides the model to backtrack to video and audio evidence, checking for factual illusions (such as fabricated images, text, or audio content that did not appear) or logical loopholes in the drafts, and correspondingly... Figure 5 The process of tracing back evidence and verifying facts.

[0075] S53, Feedback-based Decision Output: The output text of the parsing model during the validation phase. If the output contains a confirmation token, the initial answer is deemed correct and retained; if the output contains a correction token, the corrected text content is extracted as the final prediction result, achieving automatic correction of initial errors. Figure 5 The process of optimizing and outputting the answer completes the entire reflection and verification loop.

[0076] Therefore, this embodiment employs the aforementioned advertising video question-answering method based on audiovisual collaborative perception and chain verification, solving problems such as the difficulty in locating key visual cues, lack of marketing domain expertise, and unexplainable reasoning processes that are prone to illusion in existing advertising video understanding methods. This effectively improves the accuracy of advertising intent understanding and the consistency of logical reasoning. To further verify the effectiveness of this invention, this embodiment conducted a systematic comparative experiment on the publicly available advertising video question-answering benchmark dataset AdsQA. The experimental results are presented through... Figure 6 A detailed demonstration and analysis will be provided, including the specific implementation process: Figure 6 This is a rigorous accuracy comparison between this embodiment and mainstream multimodal large models on the AdsQA dataset, which is a quantitative verification of the effectiveness of this embodiment. The specific implementation and experimental design are as follows: Experimental Design Fundamentals: Test dataset: The AdsQA public advertising video Q&A benchmark dataset is used, which contains 7812 test samples. It is divided into five categories according to task type: visual perception, sentiment analysis, topic extraction, marketing strategy, and audience analysis, covering the core application scenarios of advertising video Q&A.

[0077] Evaluation Model: DeepSeek-V3 is used as a third-party impartial judge model. Relying on its excellent logical reasoning and instruction compliance capabilities, the objectivity, consistency and reproducibility of the evaluation are ensured.

[0078] Evaluation metrics: Strict accuracy is adopted. The judge model receives video metadata, standard answer and predicted answer generated by the model under test. It makes a binary judgment based on semantic consistency, completeness of key information and whether there is hallucination. The predicted answer is judged as correct only when it is semantically completely equivalent to the standard answer and does not contain any contradictory information; otherwise, it is incorrect.

[0079] Baseline comparison: Mainstream multimodal large models such as Qwen3-Base (the base model of this invention), Qwen2.5-ReAd, Qwen2.5-VL, MiniCPM-V, and LLaVA-OV were selected as baselines. All models were run in the same hardware environment (NVIDIA RTX A6000) and a unified evaluation standard was adopted to ensure the fairness of the comparison.

[0080] Key parameters: The number of oversampled advertising videos is 32, and the top-K=16 frames are selected as the key frame sequence; all input images are uniformly adjusted to 896×896 resolution; in the anchor frame generation stage, a multi-scale ratio set Scale=0.6,0.8,1.0 is set, and top-1 cropping is performed according to the density of the audiovisual-text integrated response map.

[0081] Figure 6 The results generation and analysis were based on the above experimental design. The experiment was completed under the given hardware environment (single NVIDIA RTX A6000 48GB VRAM) and software environment (Python 3.8, PyTorch framework, with expandable_segments memory optimization strategy enabled). The strict accuracy of each model in the five task types and the overall average were obtained and finally generated.

[0082] Figure 6The results show that the overall average strict accuracy of the method in this embodiment is 0.416, which is a relative improvement of 3.7% compared with the base model Qwen3-Base (0.379) without the mechanism of this invention. This proves that the audiovisual collaborative perception and chain reflection verification mechanism of this invention can effectively stimulate the reasoning potential of large models. Compared with the Qwen2.5-ReAd model (0.345) optimized for advertising scenarios, the improvement is as high as 7.1%, which verifies the core value of marketing knowledge base retrieval enhancement and structured chain verification. Among the five task types, the sentiment analysis (0.459) and topic extraction (0.455) tasks have the highest accuracy, reflecting the advantages of the embodiment of this invention in emotional atmosphere perception and core information extraction. The visual perception task is on par with Qwen3-Base (0.362), which proves that the dual-stream visual representation retains details without losing the ability to perceive the global context. Although the accuracy of the marketing strategy and audience analysis tasks is relatively slightly lower, it is still better than other comparative models, which verifies the role of domain knowledge injection in improving professional reasoning.

[0083] Experimental conclusion: This embodiment uses the AdsQA dataset as the test platform, through... Figure 4 It achieved accurate extraction of key features from advertising videos, solving the problems of asynchronous audio and video and dilution of key details; through Figure 5 It implemented domain knowledge injection and closed-loop reasoning verification, solving the problems of lack of professional knowledge and susceptibility to illusions in general models; Figure 6 Through rigorous accuracy quantification verification, the method of the embodiments of the present invention outperforms mainstream multimodal large models in the overall performance of advertising video question answering tasks, effectively improving the accuracy of advertising intent understanding, the interpretability of the reasoning process and the factual accuracy of the answer, and has strong industrial application value and generalization potential.

[0084] like Figure 2 As shown, embodiments of the present invention also provide an advertising video question-answering system based on audiovisual collaborative perception and chain verification, comprising: The data preprocessing module is used to acquire multimodal data, which includes advertising video data to be analyzed and natural language questions input by users; the multimodal data is preprocessed by extracting speech text from the video, and by oversampling the video and semantically-based keyframe optimization to obtain keyframe sequences and aligned speech text; The audiovisual collaborative anchoring module is used to construct an audiovisual and text collaborative perception and adaptive anchoring mechanism based on the obtained keyframe sequence and aligned speech text, and generate a dual-stream visual representation that includes global context and local fine perspective. The knowledge base retrieval module is used to extract semantic keywords based on the natural language questions input by the user and the generated dual-stream visual representations. At the same time, it searches in a pre-built multi-dimensional advertising knowledge base to obtain expert context information that matches the current context. The chain reasoning generation module is used to concatenate the generated dual-stream visual representation, the retrieved expert context information, and the user-input natural language question into a multimodal input sequence, so as to drive the large language model to perform chain reasoning and answer generation under structured semantic constraints and obtain the initial answer; The chain-verification optimization module is used to introduce a chain-verification mechanism based on trigger conditions, based on the initial answer. Through feedback loop, the initial answer is fact-checked and iteratively optimized to obtain the final answer.

[0085] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0086] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A question-answering method for advertising videos based on audiovisual collaborative perception and chain verification, characterized in that, The method includes: Acquire multimodal data, which includes advertising video data to be analyzed and natural language questions input by users; preprocess the multimodal data by extracting speech text from the video, and by oversampling the video and semantically-based keyframe optimization to obtain a keyframe sequence and aligned speech text; Based on the obtained keyframe sequence and aligned speech text, a co-perception and adaptive anchoring mechanism for audiovisual and textual data is constructed to generate a dual-stream visual representation that includes global context and local fine-grained perspective. Based on the natural language questions obtained from user input and the generated dual-stream visual representation, semantic keywords are extracted, and at the same time, a search is performed in a pre-built multi-dimensional advertising knowledge base to obtain expert context information that matches the current context. The generated dual-stream visual representation, the retrieved expert context information, and the user-input natural language question are concatenated into a multimodal input sequence to drive the large language model to perform chained reasoning and answer generation under structured semantic constraints, thereby obtaining the initial answer. Based on the initial answer, a chain-like verification mechanism based on trigger conditions is introduced. Through feedback loop, the initial answer is fact-checked and iteratively optimized to obtain the final answer.

2. The advertising video question-answering method based on audiovisual collaborative perception and chain verification according to claim 1, characterized in that, Acquire multimodal data, including advertising video data to be analyzed and natural language questions input by the user; preprocess the multimodal data by extracting speech text from the video, and by oversampling the video and performing semantic-based keyframe optimization to obtain a keyframe sequence and aligned speech text, including: Based on the acquired advertising video data to be analyzed, the audio text in the video is extracted. The extracted audio text is then timestamped and cleaned to obtain the aligned audio text. Based on the acquired advertising video data to be analyzed, the video is subjected to high-frequency uniform oversampling to obtain a sampled frame sequence. Pre-trained visual and language models are used to extract image features, semantic vectors of user-input natural language questions, and semantic vectors of aligned speech text from the sampled frame sequence, respectively. At the same time, the cross-modal similarity between the image features of each sampled frame and the semantic vectors of user-input natural language questions and aligned speech text is calculated. Based on the cross-modal similarity, key frames that meet a preset number are selected from the sampled frame sequence to obtain a key frame sequence.

3. The advertising video question-answering method based on audiovisual collaborative perception and chain verification according to claim 2, characterized in that, Generating the dual-stream visual representation includes: Based on the keyframe sequence, the visual space feature map of the current keyframe is extracted; at the same time, the semantic vector of the natural language question input by the user, the semantic vector of the aligned speech text, and the semantic vector of the preset OCR keywords are extracted respectively. Through cross-modal interaction, a user question heatmap, a speech heatmap, and an OCR text heatmap are generated. Based on the speech heatmap, to address the issue of audio-visual asynchrony, a time-series sliding window is set with the current keyframe as the center. The attention weights of the speech text to the current visual frame at each moment within the window are aggregated to generate a time-aligned speech heatmap. The user question heatmap and OCR text heatmap are then weighted and fused with the time-aligned speech heatmap to generate a comprehensive audiovisual and text response map. Calculate the spatial weighted average of the audiovisual and text integrated response map to locate the semantic centroid coordinates; with the centroid coordinates as the center, expand outwards in all directions in combination with the original image size and the preset multi-scale set to generate multiple candidate anchor boxes, and truncate the boundary coordinates of each candidate anchor box at the same time. Calculate the normalized response density score within each candidate anchor frame, select the anchor frame with the highest score as the final local visual region for cropping, and obtain a local fine-view image; at the same time, retain the original keyframe image after uniform scaling as the global view image; stitch the global view image and the local fine-view image together to form a two-stream visual representation that includes global context and local fine-view.

4. The advertising video question-answering method based on audiovisual collaborative perception and chain verification according to claim 3, characterized in that, Based on the natural language questions input by the user and the generated dual-stream visual representation, semantic keywords are extracted. Simultaneously, a search is performed in a pre-built multi-dimensional advertising knowledge base to obtain expert contextual information matching the current context, including: Based on the natural language questions input by the user and the dual-stream visual representation that includes global context and local fine-grained perspective, semantic keywords are extracted and mapped into text-based search query tags; Based on the search query tags, a search is conducted in a pre-constructed multidimensional advertising knowledge base to obtain multiple candidate knowledge entries that match the current context; wherein, the multidimensional advertising knowledge base includes at least the knowledge dimensions of advertising persuasion strategies, emotional stimulation mechanisms, visual metaphor mapping relationships, and target audience profile characteristics; The multiple candidate knowledge items are reordered by confidence level to select a preset number of knowledge items with the highest confidence level as expert context information.

5. The advertising video question-answering method based on audiovisual collaborative perception and chain verification according to claim 4, characterized in that, The generated dual-stream visual representation, retrieved expert context information, and user-input natural language question are concatenated into a multimodal input sequence to drive the large language model to perform chained reasoning and answer generation under structured semantic constraints, resulting in an initial answer, including: The generated dual-stream visual representation, the filtered expert context information, and the user-input natural language question are concatenated into a multimodal input sequence according to a preset format. A structured semantic constraint decoding mechanism is constructed, and a chain generation protocol of observation, analysis, reasoning, and answer is set to guide the large language model to reason and generate according to the specified cognitive path; Under the structured semantic constraint decoding mechanism, the multimodal input sequence is input into a large language model, and fine-grained semantic constraints are applied during the text generation process of the large language model. The fine-grained semantic constraints include at least the following: action refinement constraints, which are used to map general action words to specific dynamic behavior words; sentiment valence quantification constraints, which are used to extract keywords of emotional atmosphere of the scene; and visual text anchoring constraints, which are used to preferentially use OCR text information visible in the local fine-view image in the dual-stream visual representation. Based on the applied fine-grained semantic constraints, the large language model is driven to perform step-by-step reasoning according to the chain generation protocol of observation, analysis, reasoning, and answer, generating an initial answer that includes the intermediate reasoning process.

6. The advertising video question-answering method based on audiovisual collaborative perception and chain verification according to claim 5, characterized in that, Based on the initial answer, a chain-like verification mechanism based on trigger conditions is introduced. Through a feedback loop, the initial answer is fact-checked and iteratively optimized to obtain the final answer, including: Based on the generated initial answer, the system detects whether the user-input natural language question contains preset high-risk triggering keywords. If not, the initial answer is directly output as the final answer. If it is, the verification process is initiated. High-risk triggering keywords include at least keywords pointing to text details, keywords pointing to logical causality, and keywords pointing to speech content. While keeping the parameters of the large language model frozen, fact-checking prompts are constructed, and the user-input natural language question, the generated initial answer, and the generated dual-stream visual representation are input back into the large language model to guide the large language model to backtrack to multimodal evidence including the keyframe sequence and the aligned speech text, in order to evaluate whether there are factual illusions or logical loopholes in the initial answer, and generate feedback text containing confirmation tokens or correction tokens. The generated feedback text is parsed. If the feedback text contains a confirmation token, the initial answer is determined to be correct and is retained as the final answer. If the feedback text contains a correction token, the corrected answer content is extracted from the feedback text as the final answer, thus completing the iterative optimization of the initial answer.

7. An advertising video question-answering system based on audiovisual collaborative perception and chain verification, wherein the system implements the method as described in any one of claims 1 to 6, characterized in that, include: The data preprocessing module is used to acquire multimodal data, which includes advertising video data to be analyzed and natural language questions input by users; the multimodal data is preprocessed by extracting speech text from the video, and by oversampling the video and semantically-based keyframe optimization to obtain keyframe sequences and aligned speech text; The audiovisual collaborative anchoring module is used to construct an audiovisual and text collaborative perception and adaptive anchoring mechanism based on the obtained keyframe sequence and aligned speech text, and generate a dual-stream visual representation that includes global context and local fine perspective. The knowledge base retrieval module is used to extract semantic keywords based on the natural language questions input by the user and the generated dual-stream visual representations. At the same time, it searches in a pre-built multi-dimensional advertising knowledge base to obtain expert context information that matches the current context. The chain reasoning generation module is used to concatenate the generated dual-stream visual representation, the retrieved expert context information, and the user-input natural language question into a multimodal input sequence, so as to drive the large language model to perform chain reasoning and answer generation under structured semantic constraints and obtain the initial answer; The chain-verification optimization module is used to introduce a chain-verification mechanism based on trigger conditions, based on the initial answer. Through feedback loop, the initial answer is fact-checked and iteratively optimized to obtain the final answer.

Citation Information

Patent Citations

  • Method and system for solving illusion problem of large legal language model

    CN117744802A

  • Complex scene-oriented end-to-end semantic extraction system

    CN121031606A

  • Intellectual visual question and answer method and device based on big and small model collaboration and medium

    CN121350217A

  • Video question-answering system and method based on iterative multi-mode

    CN121388101A

  • Multi-modal multi-scale retrieval enhancement generation method, system and equipment applied to external knowledge questions and answers and medium

    CN121765049A

Cited By

  • Question and answer processing method, device, storage medium, and program product

    CN122198152A