Music video question answering method based on music feature guidance in few-sample scene
By using music characteristic guidance method in a small sample scenario, the time series characteristics of audio and visual information are integrated, and the knowledge advantages of large language models are combined, the multimodal data fusion problem in music video question-and-answer tasks are solved, significantly improving the generalization and reasoning ability of the model.
Patent Information
- Application Number
- CN202510257748.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
In the case of few samples, it is difficult for music video Q&A to effectively integrate audio and video modal information, resulting in limited model performance, especially when deep interactions of multimodal data are difficult to capture.
By counting the source information and extracting music characteristics, the prior knowledge of music is introduced into the multimodal fusion process, the music characteristics are used to guide the time series characteristics of the fused audio and visual information, and cross-modal representations with consistent timing are generated, and combined with the knowledge advantages of the large language model, the thinking chain prompts supplement the shortcomings of semantic information in the small sample data.
It effectively enhances the model's understanding of multimodal data, significantly improves the model's generalization and reasoning ability under data scarcity, realizes efficient inference in a small sample scenario, and has strong robustness to noise in multimodal data.
Smart Images

Figure CN120179876A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and multimodal deep learning technology, and specifically relates to a video question-answering method based on music feature guidance in a few-sample scenario. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, multimodal deep learning has gradually become one of the research hotspots. Multimodal data fusion can more comprehensively understand complex scene information by jointly modeling data of different modalities (such as images, audio, text, etc.). Among them, music videos, as a typical multimodal data, combine the characteristics of audio and video to provide users with a rich sensory experience. However, there are still many technical challenges in the analysis and understanding of music videos, especially in the case of few samples, where the limited data further exacerbates the difficulty of the task.
[0003] Currently, video question answering technology has received widespread attention in the field of multimodal research. Its goal is to generate accurate answers to questions asked by users by analyzing the video content. In the music video question answering task, it is necessary to have a deep understanding of the visual information of the video and to extract and analyze the features of the music audio. At the same time, there are complex correlations between information in different modalities. How to effectively model these correlations is a key issue in this task. This problem is particularly prominent in the scenario of few samples: insufficient data makes the model prone to overfitting, and it is difficult to capture the deep interaction between multimodal information.
[0004] Traditional video question-answering methods usually rely on large-scale annotated data and use simple feature splicing or weighted fusion methods, which makes it difficult to fully capture the deep interactive information between multiple modalities. In the scenario of few samples, the music video question-answering task faces even more severe challenges. On the one hand, the synchronization and correlation between audio and video modalities are strong, and ignoring music characteristics (such as sound source, melody, rhythm, etc.) may lead to deviations in the model's understanding of video content; on the other hand, under the condition of few samples, how to extract universal features through limited data and use music features to guide the model's spatiotemporal information fusion is a technical problem that needs to be solved urgently.
[0005] Compared with existing few-shot video question answering and few-shot music learning methods, the music video question answering task in the few-shot scenario is unique: it not only needs to process the high-dimensional features and complex associations of audio and video modalities, but also needs to fully explore the guiding role of music features in semantic understanding. For example, sound source features can be used for target positioning, which is not significant in traditional few-shot video question answering tasks. Therefore, designing a method based on music feature guidance to achieve efficient fusion of multimodal data under few-shot conditions has become an important research direction for improving the performance of music video question answering. Summary of the invention
[0006] In view of the above problems existing in the prior art, the present invention proposes a music video question answering method based on music guidance in a few-shot scenario, and constructs a few-shot scenario analysis video and audio sound source to deeply understand the music video.
[0007] The music video question answering method based on music guidance in a few-shot scenario includes the following steps:
[0008] S1. Perform few-shot setting, classify according to the types of questions and answers, and select a few samples for each type as the few-shot scenario training set.
[0009] S2. Obtain video visual information, audio information, and obtain question and answer text information, and use an encoder to encode the multi-modal information separately.
[0010] S3. Statistically collect sound source information as prior knowledge, and perform feature extraction and encoding on the sound source and its music characteristics.
[0011] S4. Use music characteristics to guide, fuse the time series features of audio and visual information, and generate temporally consistent cross-modal representations.
[0012] S5. Combine the constraints of few-shot data, utilize the knowledge advantages of the large language model, and combine music characteristics and question text to generate a chain of thought prompt for reasoning to supplement the lack of semantic information in the few-shot data.
[0013] S6. Based on the chain of thought prompt, use a spatio-temporal perception model to select time segment features and spatial region features with high relevance to the current question.
[0014] S7. Fuse the text, visual, and audio three-modal features, and use a classifier to obtain the question and answer. Calculate the loss and perform backpropagation during the training phase, and directly output the answer during the inference phase.
[0015] Specifically, to study the music video question answering method in a few-shot scenario, the present invention sets 3 different few-shot scenarios according to the modality, question, and answer categories.
[0016] Specifically, the present invention extracts visual features from the video, intercepts the video at 1 frame per second, and repeats and copies the last frame to fill the video with a length less than the maximum length. Preprocess the image, adjust the image size, adjust the image to a size of 224×224, and perform RGB channel feature normalization. Use the image encoder of the CLIP pre-trained model to extract the preprocessed image. Preprocess the text, including word segmentation and stop word removal. Use the text encoder of the CLIP pre-trained model to extract the preprocessed text. Use the VGGish model for audio feature extraction.
[0017] Specifically, the present invention conducts statistical analysis on the number and expression forms of sound sources, where the sound sources include various categories, such as background music, environmental sound effects, instrument tracks, and human voices. Key music features of the sound sources in dimensions such as timbre, pitch range, and rhythm are collected. The sound source feature encoder shares weights with the question text encoder to achieve alignment and unified encoding of cross-modal features.
[0018] Specifically, the present invention calculates the attention scores of sound sources at different times in visual and audio feature information. Based on the attention scores of the sound sources in the visual and audio modalities, the bidirectional cross-entropy loss is calculated to optimize the accuracy of the attention distribution. The audio and visual features are aligned and fused using the sound source attention scores. During this process, the attention scores are used as weights to re-weight the audio and visual features.
[0019] Specifically, the present invention includes few-shot thought chains, which contain core reasoning steps, such as time localization, spatial localization, and comparison of sound source information. The question text and the few-shot thought chain text are concatenated and input into the large language model API, and preliminary thought chain prompts are generated through its built-in reasoning ability. According to the preliminary results generated by the model and combined with the semantic features of the question text, more accurate thought chain prompts are iteratively generated. Post-processing is performed on the generated thought chain prompts to optimize their logic, semantic clarity, and matching degree to the question.
[0020] Specifically, the present invention concatenates and fuses text with visual features and combines spatio-temporal correlation. Matrix multiplication is used to achieve the three-modal fusion of the question text features and the audio-visual features. Forty-two answer categories are statistically analyzed, and a multi-layer perceptron classifier is constructed to generate the final answer. The Softmax activation function is applied to calculate the prediction probabilities of each category. During the training phase, backpropagation is performed using the cross-entropy loss and the AdamW optimizer, and during the inference phase, argmax is used to select the category with the highest prediction probability as the final answer.
[0021] The present invention proposes a method for video question answering guided by music features in a few-shot scenario. Aiming at the problems that it is difficult to effectively integrate music modality information and the model performance is limited in the few-shot scenario in the prior art, a series of innovative solutions are proposed. The present invention introduces prior knowledge of music into the multi-modal fusion process by statistically analyzing sound source information and extracting music characteristics, realizes temporal consistent modeling of audio and visual information, and effectively enhances the model's ability to understand multi-modal data. Combining the knowledge advantages of large language models, it supplements the lack of semantic information in the few-shot scenario through chain-of-thought prompting, and significantly improves the generalization ability and reasoning ability of the model under data-scarce conditions. Through a spatio-temporal perception model based on chain-of-thought prompting, the present invention can accurately select the time paragraphs and spatial region features relevant to the current question, and fuse the three-modal information to generate question-answering answers, realizing efficient reasoning in the few-shot scenario. At the same time, the model has strong robustness to noise in multi-modal data. Description of the Drawings
[0022] Figure 1 It is a flowchart of a music video question answering method guided by music features;
[0023] Figure 2 It is a flowchart of the few-shot scenario setting;
[0024] Figure 3 It is a schematic diagram of the model structure of a music video question answering method guided by music features. Detailed Embodiments
[0025] The present invention will be further described in detail below in conjunction with the drawings through specific embodiments. Similar elements in different embodiments are labeled with related similar element numbers. In the following embodiments, many details are described to make the present application better understood. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, and methods. In some cases, some operations related to the present application are not shown or described in the specification to avoid the core part of the present application being overwhelmed by excessive description. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations according to the description in the specification and general technical knowledge in the art.
[0026] In addition, the features, operations, or characteristics described in the specification can be combined in any appropriate manner to form various embodiments. At the same time, the steps or actions in the method description can also be reordered or adjusted in an obvious manner by those skilled in the art. Therefore, the various sequences in the specification and drawings are only for clearly describing a certain embodiment and do not mean that they are necessary sequences, unless it is stated that a certain sequence must be followed.
[0027] Refer to Figure 1 As shown, it is a flowchart of a few-shot scenario-based music feature-guided video question answering method according to an embodiment of the present invention. The few-shot scenario-based music feature-guided video question answering method according to an embodiment of the present invention includes the following steps:
[0028] S1. Perform few-shot setting, such as Figure 2 As shown, classify the data according to the types of questions and answers, and select a limited number of samples for each category to construct a training set in the few-shot scenario. Specifically, this step includes:
[0029] S101. For studying the music video question answering method, this embodiment selects the MUSIC-AVQA dataset. This dataset includes three types of question answering: visual, auditory, and visual-auditory fusion. The answers include 42 answers, which are: 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, more than 10, left, middle, right, saxophone, pipa, conga drum, pipa, ukulele, violin, clarinet, guzheng, piano, cello, accordion, flute, xylophone, banjo, trumpet, bagpipe, acoustic guitar, bassoon, tuba, electric bass, erhu, drum, suona. There are 5 types of question forms, which are counting, comparison, localization, existence, and temporality. Among them, visual modality question answering only includes counting and localization questions, auditory modality questions only include counting and comparison questions, and audio-visual fusion questions include all types.
[0030] S102. Set different few-shot scenarios according to modality, question, and answer categories. According to modality, question, and answer categories, set 3 different few-shot scenarios, which are 5-way k-shot, 9-way k-shot, and 42-way k-shot. One setting is 5-way k-shot, where few-shot scenarios are established for 5 types of questions respectively, and k samples are taken for each type of question. Another setting is 9-way k-shot, where few-shot scenarios are established according to the division of modality and questions, and k samples are taken for each type of question. Another setting is 42-way k-shot, where few-shot scenarios are established according to the division of 42 answer categories, and k samples are taken for each type of question.
[0031] S2. Obtain video visual information, audio information, obtain question and answer text information, and use an encoder to encode the multi-modal information separately. Specifically, this step includes:
[0032] S201: Extract visual features from the video. In this embodiment, the video is sliced into one frame per second, and the image frames corresponding to the corresponding time periods are intercepted as visual information. In order to extract 1 frame per second, the frames can be extracted according to the following rules:
[0033] F i = Frame(t i ), i = 0, 1, 2, ..., T
[0034] Among them, F i represents the frame extracted at the i-th second. For all videos, the longest video time is selected as the unified length, and for videos with insufficient duration, the last frame is copied and padded to make up the difference.
[0035] S202: Preprocess the image. Before inputting the image into the vision encoder, it is necessary to preprocess the image appropriately to ensure that the image can be effectively encoded by the model. Common image preprocessing steps: 1. Image resizing: First, the input image needs to be resized to a fixed size. For the ViT model, usually the longer side of the image is resized to 224 pixels while maintaining the aspect ratio, and the image may be cropped or padded. Then, the image is scaled to the specified input size (224×224 pixels). 2. Normalization: Normalize the RGB channels of the image to ensure that the pixel values of each channel are within the same range.
[0036] S203: To ensure that the image features can comprehensively reflect the visual content of the video, the vision encoder in the Contrastive Language-Image Pretraining (CLIP for short) is used to extract features for each frame of the image. CLIP is a multi-modal learning model that realizes cross-modal embedding space alignment through joint training of contrastive learning tasks for images and texts. CLIP can effectively understand the semantic association between images and texts. Its core idea is to learn a general representation that can perform two-way matching between images and texts through a large-scale image-text pair data. CLIP performs excellently in image classification, image-text retrieval, and generation tasks, provides efficient feature extraction capabilities for multi-modal tasks, and has significant advantages in applications in complex scenarios. In this embodiment, the text encoder and the image encoder of CLIP are respectively used as the text feature extractor and the vision feature extractor of this embodiment. The process of using CLIP to extract the visual features of a 1-frame-per-second video is mainly divided into the following steps, and the relevant formulas are as follows:
[0037]
[0038] Among them, represents the visual information of the i-th frame, and v i represents the visual feature of the i-th frame extracted by the encoder.
[0039] In the present invention, a CLIP vision encoder using a Vision Transformer (ViT) architecture is employed. The ViT model encodes images through a Transformer network, treating them as a series of image patches. Each image patch is mapped to a vector via a one-dimensional convolution and then fed into a multi-layer Transformer for self-attention mechanism calculations. The specific steps are as follows: 1. The input image is divided into blocks of a fixed size: The image is divided into blocks of 32×32 pixels. 2. Embedding layer: Each image patch is transformed into a vector of a fixed dimension (768 dimensions) through a linear mapping (usually using a fully connected layer). 3. Positional encoding: Since the Transformer does not have the ability to perceive spatial structure, positional encoding needs to be added to each image patch to ensure that the model can distinguish the spatial positions of the image patches. 4. Multi-layer self-attention: The input undergoes self-attention operations through multiple Transformer layers to model the relationships between image patches. 5. Global pooling: After self-attention calculations, a pooling operation is performed to obtain a fixed-length image feature vector.
[0040] S204: Preprocess the text, including word segmentation and stop word removal.
[0041] S205: Encode the preprocessed text using the text encoder in CLIP to generate a feature vector of the text.
[0042] F Q = CLIP_Text_Encoder(Q), i = 0, 1, 2,..., T
[0043] where Q represents the problem text information after preprocessing, and F Q represents the problem text features extracted using the encoder.
[0044] S206: Use a model based on the VGG (Visual Geometry Group) network architecture (abbreviated as VGGish) for audio feature extraction. VGGish is commonly used to extract compact and useful features from raw audio. The input of VGGish is usually a preprocessed audio signal, and the output is a fixed-length embedded feature, which is convenient for downstream tasks. Its application scenarios include audio classification, event detection, and audio encoding in multi-modal tasks. The core advantage of VGGish is that its pre-trained model can capture a wide range of audio characteristics and provide good performance in many audio analysis tasks. In this embodiment, VGGish is used as the audio feature extractor.
[0045] The VGGish model is based on the VGG architecture and pre-trained with a large amount of audio data. It can efficiently extract features such as melody, rhythm, and sound source in audio signals, generate fixed-length audio feature vectors to represent audio information.
[0046]
[0047] Among them, a i represents the audio information at time i, represents the audio feature at time i extracted by VGGish.
[0048] S3. Statistically collect sound source information as prior knowledge, and extract and encode the features of the sound source and its music characteristics. Specifically, this step includes:
[0049] S301: Statistically analyze the quantity and expression of sound sources. Sound sources include various categories, such as background music, environmental sound effects, instrument tracks, and human voices.
[0050] S302: Collect the music characteristics of different sound sources and collect their key features in dimensions such as timbre, pitch range, and rhythm.
[0051] S303: Represent the features of the sound source through the language modality so as to share weights with the question text encoder and achieve cross-modal feature alignment and unified encoding. Through this feature sharing mechanism, the information mapping and feature fusion between the audio modality and the text modality are effectively completed.
[0052] F S =CLIP_Text_Encoder(S)
[0053] Among them, S represents the sound source and its music characteristics, and F S represents the text features of the sound source and its music characteristics extracted by the encoder.
[0054] S4. Guided by music characteristics, fuse the time series features of audio and visual information to generate a temporally consistent cross-modal representation. As Figure 3 shown. Specifically, this step includes:
[0055] S401: Calculate the attention scores of the sound source at different times in the visual and audio feature information.
[0056]
[0057] Among them, a i (i = 0, 1,... T) is the audio feature information segmented from time 0 to T, and v i (i = 0, 1,... T) is the visual feature information segmented from time 0 to T, is the transposed matrix of the sound source feature information, d k is the hidden dimension, Att a is the attention score of the sound source in the audio feature information, Att v is the attention score of the sound source in the visual feature information. By performing weighted calculations on the audio and visual features respectively with the sound source feature information, an attention matrix that changes over time is generated. The calculation results respectively represent the attention distribution of the audio modality and the visual modality to different sound sources and its time variation law.
[0058] S402: Based on the attention scores of the sound source from the visual and audio modalities, calculate the bidirectional cross-entropy loss to optimize the accuracy of the attention distribution. Specifically, using the sound source feature as the target, compare the attention scores of the visual modality and the audio modality, measure the deviation in their temporal consistency, and minimize this deviation to achieve more precise feature alignment.
[0059] L av = CrossEntropy(Att a , Att v )
[0060] L va = CrossEntropy(Att v , Att a )
[0061] where L av is the cross-entropy loss from audio to visual, and L va is the cross-entropy loss from visual to audio.
[0062] S403: Use the sound source attention scores to align and fuse the audio features and visual features. During this process, take the attention scores as weights and re-weight the audio and visual features, so that the fused features can more prominently highlight the parts that are consistent with the sound source characteristics in the time dimension, thereby generating a temporally consistent cross-modal representation.
[0063]
[0064] where, A′ a is the audio feature after fusing the sound source attention and audio, and A′ v is the visual feature after fusing the sound source attention and visual.
[0065] S5: Utilize the knowledge advantage of the large language model, combine music characteristics and the problem text to generate a chain of thought prompt, and supplement the deficiency of semantic information in the few-shot data. Specifically, this step includes:
[0066] S501: Write a few-shot thought chain. The core purpose of the few-shot thought chain is to use the minimum number of examples to help the large language model understand the structure of the reasoning problem, so as to generate a reasonable reasoning process. Determine the problem type and goal, and select one or more examples similar to the target problem. The design of the few-shot examples is concise and includes core reasoning steps, such as time localization, spatial localization, sound source information comparison, etc., aiming to help the model understand how to extract key information from the input information and build a logical chain.
[0067] S502: Input the problem text and the few-shot thought chain text into the large language model, and generate a preliminary thought chain prompt through its built-in reasoning ability. The large language model parses the problem text through its pre-trained extensive corpus and context understanding ability, and combines existing knowledge to generate a draft reasoning chain that conforms to the semantics of the problem. This step does not require additional adjustment or parameter optimization of the model, and directly uses the API output to complete the construction of the preliminary prompt.
[0068] S503: According to the preliminary results generated by the model, combined with the semantic features of the problem text, iteratively generate more accurate thought chain prompts. Through multiple rounds of interactive generation and adjustment, refine the prompt content to ensure that it can cover the key semantic points in the problem and clarify the core logical path of reasoning. After each round of generation, by comparing different versions of the prompt, select the result with the best coverage and logical consistency as the intermediate product.
[0069] S504: Post-process the generated thought chain prompt to optimize its logic, semantic clarity, and matching degree with the problem. Specific methods include removing redundant information, revising inaccurate expressions, and integrating complementary content in multiple rounds of generation, and finally construct a complete and clear reasoning chain to provide reliable guiding information for the subsequent steps.
[0070] S505: Transmit the finally generated thought chain prompt to the spatio-temporal feature screening and fusion module as guiding information. Utilize the reasoning logic and problem relevance defined in the prompt, and use it as the conditional input for the subsequent model to support the efficient screening and alignment of multi-modal spatio-temporal features, ensuring that the reasoning chain can effectively serve the task goal.
[0071] S6: Based on the thought chain prompt, use the spatio-temporal perception model to select the time paragraph features and space region features highly relevant to the current problem. Specifically, this step includes:
[0072] S601: Use the thought chain prompt as the conditional input, and combine the time series and spatial features of the video to construct task context information. By parsing the thought chain prompt, extract the spatio-temporal correlation indication information therein, such as the time range involved or the spatial region of concern, to provide a guiding basis for spatio-temporal feature selection.
[0073] S602: Input the temporal and spatial features of the input video into the spatio-temporal perception model to perform a preliminary scoring on different time periods and spatial regions. Utilize the attention mechanism embedded in the spatio-temporal perception model to align the input features with the guiding information of the chain of thought prompts, and calculate the attention distribution in the temporal and spatial dimensions.
[0074] S603: According to the attention scores, screen the temporal segment features and spatial region features that are highly relevant to the current problem. Specifically, by setting a threshold or selecting several features with the highest attention scores, extract the temporal and spatial information that best fits the problem semantics, and complete the preliminary screening and extraction of features.
[0075] S604: Integrate the screened highly relevant temporal segment features and spatial region features to generate a spatio-temporally consistent feature representation. During the integration process, through feature re-weighting and normalization operations, ensure that the extracted features can accurately reflect the spatio-temporal semantic requirements of the current problem, providing a basis for subsequent multi-modal fusion steps.
[0076] S7: Fuse the text, visual, and audio multi-modal features, and use a classifier to obtain the Q&A answer. Calculate the loss and perform backpropagation during the training phase, and directly output the answer during the inference phase. Specifically, this step includes:
[0077] S701: Fuse the text and visual features and combine spatio-temporal correlations. Connect the temporally highly relevant audio features, temporally highly relevant visual features, and spatially highly relevant visual features extracted in step S3. Subsequently, use a fully connected layer to transform the concatenated multi-modal features and project them into the same feature space as the text features for unified representation.
[0078] S702: Achieve the three-modal fusion of the question text features and the audio-visual features. The specific method is to perform an element-wise multiplication operation on the audio-visual features extracted in the previous step and the question text features to complete the feature fusion.
[0079] S703: Count the question categories and answer types, and construct a classifier to generate the final answer. Collect all the questions and answer categories in the training set, count the number of answer categories and perform one-hot encoding. The classifier first uses a linear layer with weights \(W\in R\) D ×C , where \(D\) is the dimension of the fusion layer and \(C\) is the number of answer categories.
[0080] S704: Apply the Softmax activation function to calculate the prediction probabilities of each category. Use cross-entropy to calculate the loss between the prediction probabilities and the true answers during the training phase, and perform backpropagation using the AdamW optimizer. Select the category with the highest prediction probability as the answer output by the model during the inference phase.
[0081] The above uses specific examples to elaborate on the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the art to which the present invention pertains, based on the idea of the present invention, several simple deductions, deformations or substitutions can also be made.
Claims
1. A music video question answering method based on music guidance in a few-sample scenario, characterized in that: The method comprises the following steps: S1. Perform a few-sample setting, classify questions and answers into categories, and select a few samples from each category as the few-sample scenario training set; S2, obtaining video visual information, audio information, and question and answer text information, and using an encoder to encode the multimodal information separately; S3, counting and collecting sound source information as prior knowledge, and performing feature extraction and encoding of the sound source and its musical characteristics; S4, using music characteristics to guide, fuse the time series features of audio and visual information to generate temporally consistent cross-modal representations; S5. Combined with the constraints of few sample data, using the knowledge advantages of the large language model, combined with the characteristics of music and question text, generate thought chain prompts for reasoning to supplement the lack of semantic information in the few sample data; S6. Based on the thought chain prompts, use the spatiotemporal perception model to select time segment features and spatial region features that are highly relevant to the current problem; S7. Integrate the three-modal features of text, vision, and audio, and use the classifier to obtain the answer to the question. Calculate the loss and backpropagate in the training phase, and directly output the answer in the inference phase.
2. The music video question answering method based on music guidance in a few-sample scenario as claimed in claim 1, characterized in that: The step 1 includes: setting three different few-sample scenarios according to the modality, question and answer categories, namely 5-way k-samples, 9-way k-samples and 42-way k-samples.
3. The music video question answering method based on music guidance in a few-sample scenario as claimed in claim 1, characterized in that: The step 2 comprises: S201: Extract visual features from video; S202: preprocessing the image, adjusting the image size, and normalizing the RGB channels of the image; S203: Using the image encoder of the CLIP pre-trained model to extract features from the pre-processed image; S204: preprocessing the text, including word segmentation and removal of stop words; S205: Encode the preprocessed text using a text encoder of the CLIP pre-trained model to generate a feature vector of the text; S206: Use the VGGish model to extract audio features.
4. The music video question-answering method based on music guidance in a few-sample scenario as claimed in claim 1, characterized in that: The step 3 comprises: Conduct statistical analysis on the number of sound sources and their expressions, which include various categories; Collect key musical features of the sound source in terms of timbre, range, and rhythm; The sound source feature encoder and the question text encoder share weights to achieve cross-modal feature alignment and unified encoding.
5. The music video question answering method based on music guidance in a few-sample scenario as claimed in claim 1, characterized in that: The step 4 comprises: Calculate the attention score of the sound source at different times in the visual and audio feature information; Based on the attention scores of the visual and audio modalities to the sound source, a bidirectional cross entropy loss is calculated to optimize the accuracy of the attention distribution; The audio and visual features are aligned and fused using the sound source attention score. In this process, the attention score is used as the weight to reweight the audio and visual features.
6. The music video question answering method based on music guidance in a few-sample scenario as claimed in claim 1, characterized in that: The step 5 comprises: Write a few-sample thinking chain, identify the question type and goal, and include the core reasoning steps; The question text and a few sample thought chain texts are spliced and input into the large language model, and the initial thought chain prompts are generated through its built-in reasoning ability; Based on the preliminary results generated by the model and combined with the semantic features of the question text, more accurate thought chain prompts are iteratively generated; Post-process the generated thought chain prompts to optimize their logic, semantic clarity and matching degree to the problem; The final generated thought chain prompts are passed to the spatiotemporal feature screening and fusion module as guidance information.
7. The music video question answering method based on music guidance in a few-sample scenario as claimed in claim 1, characterized in that: The step 6 comprises: The thought chain prompt is used as a conditional input, and the temporal sequence and spatial features of the video are combined to construct task context information. Input the time series features and spatial features of the video into the spatiotemporal perception model, and make preliminary scores for different time periods and spatial regions; According to the attention score, the time segment features and spatial region features that are highly relevant to the current problem are screened; The filtered high-correlation time segment features are integrated with the spatial region features to generate a feature representation with temporal and spatial consistency.
8. The music video question answering method based on music guidance in a few-sample scenario as claimed in claim 1, characterized in that: The step 7 comprises: Fusion of text and visual features, combined with spatiotemporal correlation; Matrix multiplication is used to achieve trimodal fusion of question text features and audio and visual features; Count the types of questions and answers, and build a multi-layer perceptron classifier to generate the final answer; The Softmax activation function is applied to calculate the predicted probability of each category. The cross entropy loss and AdamW optimizer are used for back propagation in the training phase, and argmax is used in the inference phase to select the category with the highest predicted probability as the final answer.
Citation Information
Cited By
Self-adaptive three-dimensional large language model system based on query guidance
CN120849595A
Large language model sentiment analysis method based on multi-modal thinking chain alignment
CN121980314A
Knowledge graph question and answer reasoning method based on potential energy increment triggering
CN122154952A